EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
v0.dev Template CTC Disk Qwen2 Python Password Firewall Statistics Land PIP Interview Cloudreve Food Ubuntu YOLO 腾讯云 Tiktoken 证件照 云服务器 Data tar ResNet-50 C++ LoRA Nginx git Safetensors CC TensorFlow Llama 搞笑 DeepStream FP16 OpenCV Random uWSGI Hotel SVR VPN Excel Gemma Input Clash Breakpoint LLM Mixtral Vim Git 算法题 财报 Vmess mmap 论文速读 EXCEL FP32 Docker Permission 顶会 Quantization VGG-16 版权 Image2Text FP64 Attention SQL GoogLeNet CAM Pytorch PDF Proxy Math Plate Bert CLAP Website WebCrawler 签证 Transformers 图标 Rebuttal GIT 阿里云 多进程 GPTQ 继承 Use Conda Linux NLP FP8 ONNX JSON SQLite HaggingFace diffusers TSV Claude AI OCR PyTorch Hungarian Video WAN HuggingFace Paper Michelin CV Windows Pickle Freesound COCO Translation Quantize 多线程 LaTeX Tracking TTS v2ray UNIX FastAPI Jetson 域名 LeetCode Numpy UI Knowledge Django torchinfo ChatGPT transformers VSCode Color uwsgi Hilton Qwen RL Plotly PyCharm DeepSeek Zip ModelScope API logger BeautifulSoup XGBoost 论文 Domain Pandas Diagram llama.cpp Web OpenAI Sklearn Miniforge hf Algorithm Streamlit BTC Anaconda IndexTTS2 Tensor Paddle Crawler LLAMA GGML Bipartite Bin Animate Bitcoin SPIE 净利润 PDB Shortcut CEIR Magnet 飞书 Logo 音频 公式 Heatmap scipy 第一性原理 printf Review Github Pillow Markdown 强化学习 InvalidArgumentError 递归学习法 icon Agent Dataset News Search CSV Ptyhon BF16 tqdm NLTK GPT4 ms-swift TensorRT Augmentation MD5 关于博主 Baidu SAM RGB Card NameSilo Jupyter Distillation 图形思考法 git-lfs Google CUDA XML 报税 FlashAttention Datetime QWEN Qwen2.5 RAR Base64
站点统计

本站现有博文333篇,共被浏览917712

本站已经建立2622天!

热门文章
文章归档
回到顶部