EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
Vim Permission Color Mixtral Interview LLAMA Food CEIR Pytorch 多进程 Streamlit Tracking llama.cpp VPN Markdown Windows OpenAI Clash ModelScope Proxy Tensor Firewall 递归学习法 CUDA Logo YOLO 公式 Agent printf WAN 图标 QWEN FP8 OCR ms-swift LoRA Augmentation Breakpoint UNIX 云服务器 TSV VSCode Disk Attention 飞书 Search SQLite DeepStream MD5 Datetime git HuggingFace git-lfs hf tar News Transformers Math Base64 v2ray ChatGPT Paddle Zip FastAPI Excel Quantize logger CV COCO Linux 腾讯云 Translation Plotly 继承 SQL WebCrawler GIT ResNet-50 FlashAttention PyTorch Hotel XML 证件照 PyCharm Hungarian Qwen2.5 GPTQ CSV Diagram Template Google CAM RGB Heatmap mmap Git BF16 TTS transformers Crawler Magnet Miniforge 论文速读 Pillow 阿里云 Card 域名 净利润 EXCEL Tiktoken C++ SVR NLTK 版权 XGBoost AI Shortcut LeetCode Web Michelin Data 签证 Plate Review Distillation Jetson 关于博主 Ptyhon Llama LLM Docker VGG-16 NameSilo SAM 多线程 RAR RL FP16 FP32 uwsgi TensorRT Video Python Bipartite Github Numpy Hilton BTC GGML 报税 音频 Safetensors PDF CC CLAP Ubuntu HaggingFace Nginx Pandas UI Input 图形思考法 Knowledge 顶会 JSON SPIE 算法题 v0.dev Bin Django Dataset Land Quantization Password Conda PIP CTC Baidu uWSGI ONNX Image2Text Domain Use IndexTTS2 Gemma Rebuttal diffusers Website BeautifulSoup Anaconda Sklearn torchinfo Paper 论文 Statistics LaTeX FP64 强化学习 Cloudreve Vmess TensorFlow Bert Freesound NLP GoogLeNet 搞笑 PDB OpenCV API Pickle Random Claude Animate Algorithm DeepSeek tqdm Bitcoin Jupyter Qwen2 财报 Qwen scipy 第一性原理 InvalidArgumentError GPT4 icon
站点统计

本站现有博文332篇,共被浏览900580

本站已经建立2602天!

热门文章
文章归档
回到顶部