EADST

llama.cpp: Efficient 6-bit Data Packing in an 8-bit Array

This code snippet, adapted from llama.cpp by ggerganov, demonstrates a method for efficiently packing 6-bit values into an 8-bit uint8 array. It involves scaling, clamping, and bitwise manipulation to optimize or compress data, suitable for specific processing or hardware requirements.

// Initialize inverse scale factor with a fixed scaling offset and the maximum scale value.
float iscale = -32.f/max_scale;
// QK_K = 256. Iterate over a subset of the scales array, determined by QK_K divided by 16.
for (int j = 0; j < QK_K/16; ++j) {
    // Scale and round the j-th element of the scales array to the nearest integer.
    int8_t l = nearest_int(iscale * scales[j]);

    // Clamp the value of l to the range [-32, 31] and normalize it to [0, 63].
    l = MAX(-32, MIN(31, l)) + 32;

    // Store the 0-7th scale lower 4 bits of l in y[i].scales if in the first half of the loop.
    if (j < 8) {
        y[i].scales[j] = l & 0xF;
    } 
    // In the second half, store the 8-15th scale lower 4 bits of l into the higher 4 bits of y[i].scales at j-8.
    else {
        y[i].scales[j-8] |= ((l & 0xF) << 4);
    }

    // Shift the higher 4 bits of l to the lower positions.
    l >>= 4;

    // Calculate the index for storing the lower 2 bits(previous l 2 higher bits) of the shifted l and store them in y[i].scales.
    // The specific position in the array is determined by a combination of modulo and division operations.
    y[i].scales[j % 4 + 8] |= (l << (2 * (j / 4)));
}

The key aspects of this code include:

  • Scaling and Normalization: Adjusts the data values to a suitable range for bit manipulation.
  • Bitwise Operations: Utilizes masking (&), shifting (<<, >>), and bitwise OR (|=) to pack data efficiently.
  • Data Optimization: The method packs data into a smaller space, allowing for efficient use of memory and potentially faster processing.

This approach is particularly useful in scenarios where memory optimization is crucial, such as in embedded systems or when dealing with large datasets.

相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
AI MD5 Translation icon CAM Augmentation tqdm Docker TTS v2ray Google Quantize Breakpoint C++ Sklearn FP8 Plotly Algorithm CUDA 递归学习法 YOLO 顶会 Qwen2.5 Permission SVR Tensor Bitcoin Conda 图标 Paper COCO 飞书 Ptyhon InvalidArgumentError XGBoost TensorFlow 多进程 Logo Distillation NLTK Quantization Nginx Review torchinfo GIT VPN Numpy Pillow OCR Math ms-swift CV CLAP PyCharm Firewall Qwen2 Vmess 报税 ONNX CC Disk Use 版权 Pytorch mmap Transformers Github API网关 TSV BeautifulSoup Attention VSCode EXCEL QWEN Streamlit Ubuntu NLP Django 音频 Agent Pickle Land Proxy Anaconda Website 强化学习 RAR NameSilo Plate Bert Heatmap BF16 Git Gemma Safetensors RGB Paddle Domain Crawler Interview Statistics Cloudreve uwsgi HuggingFace Magnet llama.cpp PIP Harness GoogLeNet 域名 Hungarian 财报 PDF OpenAI git UNIX Claude Freesound v0.dev CTC XML ModelScope RL 图形思考法 TensorRT Windows 算法题 DeepStream Llama 论文 logger 签证 证件照 Pandas Search FP32 ResNet-50 净利润 Rebuttal Michelin HaggingFace git-lfs Linux GPTQ Dataset GPT4 WebCrawler FP64 LaTeX Card Color WAN SQL 第一性原理 JSON Web API CEIR LoRA Template Excel PDB Base64 uWSGI News Clash 关于博主 Data Diagram 云服务器 PyTorch SPIE Tracking ChatGPT Hilton Food LeetCode Qwen Bin SQLite diffusers tar GGML IndexTTS2 Input hf 阿里云 多线程 FP16 Baidu printf UI Jupyter 继承 Markdown Shortcut Password Random scipy SAM Hotel VGG-16 Python FastAPI Mixtral Datetime Zip 公式 BTC FlashAttention 腾讯云 OpenCV Animate Tiktoken transformers LLM Bipartite Video DeepSeek Knowledge 论文速读 Jetson Miniforge Image2Text LLAMA Vim CSV 搞笑
站点统计

本站现有博文336篇,共被浏览940735

本站已经建立2651天!

热门文章
文章归档
回到顶部