EADST

Quick Review: ZeroQuant-FP

ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Highlights:

  • FP4 Weight Quantization: Implements 4-bit floating-point (FP4) quantization for model weights.
  • FP8 Activation Quantization: Utilizes 8-bit floating-point (FP8) quantization for activations, optimizing the balance between performance and precision.
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
Quantization FP16 Tiktoken Sklearn CC 关于博主 Logo Quantize Plate Translation Crawler ResNet-50 Harness 音频 Mixtral Rebuttal Food Vmess 强化学习 FP8 EXCEL BF16 CTC 算法题 OCR NLTK TensorFlow TensorRT Hungarian 第一性原理 Qwen2.5 Use Claude Password CV LLAMA Github 递归学习法 XGBoost BTC Paddle Algorithm Permission 财报 净利润 COCO Firewall Card ms-swift 腾讯云 公式 Qwen2 VSCode PDF TTS git Hilton Excel XML Attention Jupyter v2ray logger SQL Template Ubuntu Math FP64 阿里云 GPTQ Michelin scipy 图形思考法 Numpy Video MD5 Statistics Safetensors 版权 Streamlit Tracking NLP CEIR PyTorch RAR torchinfo 论文速读 Transformers diffusers transformers Windows PDB 飞书 API网关 Knowledge NameSilo ModelScope mmap 证件照 Shortcut Baidu RGB 云服务器 搞笑 图标 Random Pillow Data OpenCV Clash Web 论文 RL SAM Ptyhon PIP Jev 继承 Markdown Heatmap Search Python DeepSeek Bert Input BeautifulSoup Anaconda 顶会 Linux Proxy Gemma SVR Bitcoin Breakpoint Qwen 域名 Base64 FlashAttention 报税 QWEN GGML Image2Text CLAP Bin FP32 Nginx Pytorch ChatGPT HuggingFace Color Magnet Distillation CUDA Agent GPT4 JSON UI Diagram Bipartite Dataset Animate CAM Git Interview Freesound SQLite VGG-16 WAN CSV Augmentation Llama GoogLeNet uWSGI FastAPI UNIX Jetson 签证 Pandas Datetime Cloudreve Docker Tensor SPIE Land 多进程 icon InvalidArgumentError API v0.dev OpenAI Google Conda tqdm Zip llama.cpp Website hf DeepStream LoRA Plotly Paper LLM News AI git-lfs Miniforge PyCharm Django Hotel TSV WebCrawler GIT tar ONNX VPN Pickle Disk uwsgi LaTeX LeetCode Domain printf YOLO Review Vim IndexTTS2 C++ 多线程 HaggingFace
站点统计

本站现有博文337篇,共被浏览955756次

本站已经建立2666天!

热门文章
文章归档
回到顶部