Hugging Face Blog·· 2023-05-24精选AI 评分88
Hugging Face 整合 4-bit 量化与 QLoRA 技术
Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
AI 导读
Hugging Face 在 transformers 中整合了 bitsandbytes 的 4-bit 量化与 QLoRA 技术,支持在消费级 GPU 上运行和微调大模型。QLoRA 通过 4-bit NormalFloat 存储和双重量化降低显存占用,使单卡微调 65B 参数模型成为可能。
推荐理由
官方整合了4-bit量化与QLoRA技术,提供了在消费级硬件上运行和微调大模型的完整工具链。
来源:Hugging Face Blog · huggingface.co