EADST

llama.cpp: Efficient 6-bit Data Packing in an 8-bit Array

This code snippet, adapted from llama.cpp by ggerganov, demonstrates a method for efficiently packing 6-bit values into an 8-bit uint8 array. It involves scaling, clamping, and bitwise manipulation to optimize or compress data, suitable for specific processing or hardware requirements.

// Initialize inverse scale factor with a fixed scaling offset and the maximum scale value.
float iscale = -32.f/max_scale;
// QK_K = 256. Iterate over a subset of the scales array, determined by QK_K divided by 16.
for (int j = 0; j < QK_K/16; ++j) {
    // Scale and round the j-th element of the scales array to the nearest integer.
    int8_t l = nearest_int(iscale * scales[j]);

    // Clamp the value of l to the range [-32, 31] and normalize it to [0, 63].
    l = MAX(-32, MIN(31, l)) + 32;

    // Store the 0-7th scale lower 4 bits of l in y[i].scales if in the first half of the loop.
    if (j < 8) {
        y[i].scales[j] = l & 0xF;
    } 
    // In the second half, store the 8-15th scale lower 4 bits of l into the higher 4 bits of y[i].scales at j-8.
    else {
        y[i].scales[j-8] |= ((l & 0xF) << 4);
    }

    // Shift the higher 4 bits of l to the lower positions.
    l >>= 4;

    // Calculate the index for storing the lower 2 bits(previous l 2 higher bits) of the shifted l and store them in y[i].scales.
    // The specific position in the array is determined by a combination of modulo and division operations.
    y[i].scales[j % 4 + 8] |= (l << (2 * (j / 4)));
}

The key aspects of this code include:

  • Scaling and Normalization: Adjusts the data values to a suitable range for bit manipulation.
  • Bitwise Operations: Utilizes masking (&), shifting (<<, >>), and bitwise OR (|=) to pack data efficiently.
  • Data Optimization: The method packs data into a smaller space, allowing for efficient use of memory and potentially faster processing.

This approach is particularly useful in scenarios where memory optimization is crucial, such as in embedded systems or when dealing with large datasets.

相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
WAN DeepSeek RL Tiktoken Food OpenAI Shortcut 音频 Ubuntu LeetCode tar Tensor SAM Paddle YOLO Bitcoin Search git Transformers Michelin FastAPI Safetensors Plate 签证 Baidu Mixtral 飞书 FP64 关于博主 EXCEL Algorithm Bert Qwen Web Bin Card Knowledge 多线程 图标 PyTorch 多进程 Pillow GPTQ ChatGPT BeautifulSoup Breakpoint FP16 版权 OCR Google Land PDF GIT Vim Django Firewall XML UI Zip Disk ModelScope Harness Gemma 论文 Crawler 腾讯云 VGG-16 Distillation NLP C++ PDB Logo Pytorch Python Bipartite 继承 DeepStream CV Ptyhon Base64 Random API Jupyter Windows Git Input RGB Github OpenCV Tracking Domain 净利润 scipy Animate CSV Clash 强化学习 Website SQL 云服务器 证件照 Streamlit Numpy WebCrawler CTC CAM MD5 tqdm VSCode SPIE SVR Markdown AI Math Magnet PyCharm Qwen2.5 BTC IndexTTS2 API网关 Proxy Hotel Template hf uwsgi Pandas XGBoost GoogLeNet HuggingFace 论文速读 Excel Statistics LoRA Hungarian Llama VPN CLAP mmap 财报 Miniforge 图形思考法 Color Agent 公式 LLAMA Heatmap CC LLM FP8 算法题 Nginx Quantize 递归学习法 Plotly News 顶会 Permission transformers Video v0.dev UNIX BF16 Rebuttal GPT4 Qwen2 RAR FP32 uWSGI 报税 Diagram Jetson Vmess git-lfs CUDA NameSilo Image2Text HaggingFace Password logger JSON Claude QWEN TSV Translation diffusers InvalidArgumentError icon torchinfo 搞笑 TensorFlow NLTK Pickle Sklearn Jev llama.cpp 阿里云 Hilton printf Conda 第一性原理 ResNet-50 GGML v2ray SQLite 域名 LaTeX Linux COCO ONNX TTS CEIR FlashAttention Datetime Freesound ms-swift Augmentation TensorRT Use Data Anaconda Quantization Review Cloudreve Attention PIP Paper Dataset Interview Docker
站点统计

本站现有博文337篇,共被浏览964395次

本站已经建立2676天!

热门文章
文章归档
回到顶部