EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
TensorRT 公式 Streamlit CTC FlashAttention FastAPI Diagram VSCode Jetson git-lfs Vmess CEIR Website 顶会 证件照 Qwen2.5 音频 llama.cpp Card WebCrawler Animate Pillow NLTK GIT Bert Conda logger Llama PDB Quantization Markdown SQL CC Crawler 第一性原理 GPTQ ModelScope Clash 图形思考法 TTS MD5 YOLO 论文速读 多线程 FP64 Linux RAR Pickle EXCEL Google Tensor 递归学习法 算法题 QWEN BF16 Statistics 域名 Bitcoin OCR OpenAI Mixtral Tracking Random scipy tqdm Augmentation Tiktoken Claude COCO Miniforge v0.dev CV LaTeX Baidu Numpy 强化学习 Food Zip UNIX SAM Data Base64 HaggingFace WAN Hungarian Bipartite Jupyter CLAP 腾讯云 Attention Qwen2 Freesound transformers Web Vim Ptyhon RL Ubuntu diffusers Bin Image2Text Disk uWSGI IndexTTS2 LLM Interview Breakpoint Windows Github Proxy Rebuttal Template ResNet-50 DeepSeek mmap 财报 报税 CUDA OpenCV Algorithm GoogLeNet NameSilo News Nginx 图标 Excel Math Harness hf VGG-16 PyTorch v2ray Pandas torchinfo Permission CAM VPN BTC API InvalidArgumentError 版权 Logo SVR 签证 多进程 BeautifulSoup 飞书 printf Password LeetCode JSON LLAMA ms-swift Cloudreve Plate icon 关于博主 LoRA Quantize RGB Firewall XGBoost Hilton Python uwsgi TensorFlow tar Transformers Pytorch 净利润 论文 CSV Docker TSV FP8 Safetensors UI Magnet Paper SQLite Land C++ 云服务器 FP16 HuggingFace Translation 继承 Distillation ChatGPT 搞笑 Anaconda FP32 Datetime DeepStream Git Agent Heatmap Qwen Plotly API网关 NLP Knowledge XML 阿里云 PIP Sklearn SPIE Shortcut AI Search Video Jev Input ONNX Michelin Dataset git Color PyCharm Domain PDF Paddle Review Hotel GPT4 Gemma Django GGML Use
站点统计

本站现有博文337篇,共被浏览957717次

本站已经建立2669天!

热门文章
文章归档
回到顶部