EADST

Quick Review: ZeroQuant-FP

ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Highlights:

  • FP4 Weight Quantization: Implements 4-bit floating-point (FP4) quantization for model weights.
  • FP8 Activation Quantization: Utilizes 8-bit floating-point (FP8) quantization for activations, optimizing the balance between performance and precision.
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
mmap Shortcut CSV Tensor Diagram UNIX API TTS scipy RAR BF16 VSCode Land 签证 OpenCV Sklearn CUDA OCR 论文速读 Streamlit News hf PDF InvalidArgumentError 顶会 Freesound TSV Jupyter 图标 SVR Zip Firewall BTC uwsgi GGML FP64 PDB Transformers 飞书 Website Paddle DeepSeek Food Linux 递归学习法 CLAP 第一性原理 图形思考法 Docker 多线程 Permission DeepStream Quantize Distillation SQLite LeetCode Pillow Paper v2ray CTC NLTK FP8 git Ptyhon SAM Breakpoint icon RL Use SPIE 继承 C++ VPN transformers Safetensors Input Bert Michelin Numpy Bitcoin FP16 GIT Knowledge Magnet Clash Pytorch LoRA 腾讯云 Google Excel Heatmap Random Qwen2.5 RGB Quantization Miniforge ChatGPT Augmentation GPT4 版权 搞笑 git-lfs Domain Git GoogLeNet COCO WAN TensorFlow tqdm Attention ResNet-50 Card 论文 HuggingFace UI 证件照 FlashAttention Django Template CEIR Password Disk 报税 Animate Datetime Qwen Base64 GPTQ uWSGI 净利润 Plotly Tiktoken PIP Mixtral 财报 Color v0.dev diffusers Tracking Logo PyTorch Algorithm 域名 CAM Anaconda Search VGG-16 公式 Claude Conda OpenAI XML 多进程 YOLO NameSilo llama.cpp FP32 Rebuttal Bin QWEN printf tar Llama Hotel Hilton Pickle Video Web Translation 音频 AI Jetson Windows torchinfo Dataset IndexTTS2 Pandas Vmess SQL JSON 云服务器 Github CC Image2Text Bipartite 算法题 Data HaggingFace TensorRT Math Hungarian WebCrawler Statistics 关于博主 Markdown EXCEL FastAPI ms-swift Proxy Vim Baidu Qwen2 Agent 强化学习 Crawler NLP LLM Review ModelScope 阿里云 PyCharm LaTeX Gemma Cloudreve Ubuntu Python logger Nginx BeautifulSoup LLAMA ONNX Interview Plate MD5 XGBoost CV
站点统计

本站现有博文334篇,共被浏览930806

本站已经建立2639天!

热门文章
文章归档
回到顶部