Hugging Face Blog·· 2024-05-16精选AI 评分70
Hugging Face Transformers 支持 KV Cache 量化以优化长文本生成
Unlocking Longer Generation with Key-Value Cache Quantization
AI 导读
Hugging Face 在 Transformers 库中集成 KV Cache 量化功能,支持 int2 至 int4 精度,可显著降低长上下文生成的显存占用。实验显示 int4 精度在保持模型质量的同时,相比 fp16 可实现约 2.5 倍的显存节省,在 80GB A100 上支持最高 128k tokens 上下文。
推荐理由
官方在 Transformers 中集成 KV Cache 量化功能,提供 int2 至 int4 精度选项,显著降低长文本生成的显存占用。
来源:Hugging Face Blog · huggingface.co