EADST

Sharding and SafeTensors in Hugging Face Transformers

In the Hugging Face transformers library, managing large models efficiently is crucial, especially when working with limited disk space or specific file size requirements. Two key features that help with this are sharding and the use of SafeTensors.

Sharding

Sharding is the process of splitting a large model's weights into smaller files or "shards." This is particularly useful when dealing with large models that exceed file size limits or when you want to manage storage more effectively.

Usage

To shard a model during the saving process, you can use the max_shard_size parameter in the save_pretrained method. Here's an example:

# Save the model with sharding, setting the maximum shard size to 1GB
model.save_pretrained('./model_directory', max_shard_size="1GB")

In this example, the model's weights will be divided into multiple files, each not exceeding 1GB. This can make storage and transfer more manageable, especially when dealing with large-scale models.

SafeTensors

The safetensors library provides a new format for storing tensors in a safe and efficient way. Unlike traditional formats like PyTorch's .pt files, SafeTensors ensures that the tensor data cannot be accidentally executed as code, offering an additional layer of security. This is particularly important when sharing models across different systems or with the community.

Usage

To save a model using SafeTensors, simply specify the safe_serialization parameter when saving:

# Save the model using SafeTensors format
model.save_pretrained('./model_directory', safe_serialization=True)

This will create files with the .safetensors extension, ensuring the saved tensors are stored safely.

Combining Sharding and SafeTensors

You can combine both sharding and SafeTensors to save a large model securely and efficiently:

# Save the model with sharding and SafeTensors
model.save_pretrained('./model_directory', max_shard_size="1GB", safe_serialization=True)

This setup splits the model into shards, each in the SafeTensors format, offering both manageability and security.

Conclusion

By leveraging sharding and SafeTensors, Hugging Face transformers users can handle large models more effectively. Sharding helps manage file sizes, while SafeTensors ensures the safe storage of tensor data. These features are essential for anyone working with large-scale models, providing both practical and security benefits.

相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
顶会 Interview API Streamlit VSCode Michelin git-lfs ms-swift Review Animate JSON Django 强化学习 Algorithm Template QWEN llama.cpp Pytorch Paddle Vmess TensorFlow ONNX 公式 NameSilo 多线程 XML Google Attention 腾讯云 WAN Disk Crawler Agent tqdm Website Pillow Hilton Safetensors Hotel MD5 BTC NLTK 证件照 hf Qwen2.5 签证 GGML LLM Bert Datetime Windows RL Logo VPN Search Sklearn uWSGI Math Data 版权 Vim diffusers 音频 CEIR Hungarian 图标 Food Paper Numpy AI InvalidArgumentError PDB 图形思考法 FastAPI DeepStream TensorRT transformers Bin Distillation Color Land Python logger PyCharm LeetCode Translation SVR Clash LoRA FlashAttention scipy Cloudreve XGBoost OpenCV Video Magnet FP32 CSV 云服务器 OCR tar ModelScope Conda printf 报税 继承 News Claude C++ FP64 Gemma 递归学习法 Tracking Statistics icon Jupyter Git VGG-16 HuggingFace Proxy BeautifulSoup NLP DeepSeek 财报 Web Mixtral Markdown IndexTTS2 Input PDF GIT GPT4 GPTQ 搞笑 v2ray Breakpoint Freesound UI Knowledge 飞书 WebCrawler SQLite 域名 多进程 论文 Domain BF16 YOLO CTC OpenAI Rebuttal Use TSV 算法题 Base64 Firewall Random PyTorch Card RGB Quantize Heatmap HaggingFace CV 关于博主 Qwen Docker Github EXCEL Ptyhon Anaconda SAM Pickle 第一性原理 FP8 Nginx Transformers Pandas torchinfo 阿里云 RAR Shortcut Ubuntu uwsgi Diagram Jetson Password LaTeX Augmentation UNIX CC Image2Text Linux Permission ChatGPT Plotly Zip CAM mmap COCO CUDA LLAMA Llama Tiktoken Dataset Qwen2 论文速读 Plate Bipartite Miniforge SQL v0.dev GoogLeNet FP16 Bitcoin CLAP git 净利润 Baidu PIP TTS Tensor Quantization ResNet-50 SPIE Excel
站点统计

本站现有博文333篇,共被浏览921348

本站已经建立2627天!

热门文章
文章归档
回到顶部