EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
第一性原理 Rebuttal Logo Docker PyTorch JSON Animate Qwen2 LLM Bin API InvalidArgumentError uwsgi Streamlit Disk 腾讯云 Vmess 报税 News 强化学习 QWEN Statistics Magnet 飞书 OCR PDF GPTQ VPN Crawler 顶会 Miniforge Plate GoogLeNet Diagram Sklearn Ubuntu Baidu Tiktoken 递归学习法 财报 OpenCV Review NLTK CLAP Tensor CTC RAR Input Data Claude CAM DeepSeek Hotel 域名 净利润 CSV Food OpenAI Proxy Interview Pytorch Domain Paddle mmap VGG-16 Cloudreve 云服务器 Color SPIE RGB Pandas Firewall Bert FlashAttention Augmentation RL Breakpoint tar Mixtral Gemma Safetensors BF16 FP16 WebCrawler LLAMA Attention Linux Land DeepStream HuggingFace printf IndexTTS2 Password Clash Datetime Nginx HaggingFace BTC Paper CC Random 继承 PyCharm TensorFlow Zip 论文 git-lfs BeautifulSoup EXCEL scipy Dataset Windows XGBoost GIT Plotly Transformers Quantize FP8 Translation 关于博主 TSV Vim CUDA Knowledge Pickle Algorithm UI 阿里云 COCO Ptyhon transformers Markdown NameSilo FP64 Base64 Template Use SQL LaTeX torchinfo WAN Shortcut Excel XML 图标 ModelScope SQLite 签证 v2ray PIP Llama Distillation Conda Agent ChatGPT Michelin NLP FP32 音频 Jetson logger GPT4 公式 Google CEIR Permission Hungarian Anaconda YOLO Jupyter tqdm Search Hilton llama.cpp Qwen2.5 SVR Math LoRA FastAPI Heatmap Python hf ms-swift 版权 MD5 Bitcoin icon v0.dev PDB 搞笑 TTS 多进程 ONNX UNIX Git Website Freesound Tracking ResNet-50 uWSGI Pillow Django AI TensorRT VSCode diffusers 证件照 图形思考法 C++ git Image2Text Bipartite LeetCode Numpy Web 论文速读 GGML Github CV Qwen Quantization Video 多线程 算法题 Card SAM
站点统计

本站现有博文334篇,共被浏览935893

本站已经建立2646天!

热门文章
文章归档
回到顶部