跳到正文
原文
Hugging Face Blog·· 2024-05-16精选AI 评分70

Hugging Face Transformers 支持 KV Cache 量化以优化长文本生成

Unlocking Longer Generation with Key-Value Cache Quantization

AI 导读

Hugging Face 在 Transformers 库中集成 KV Cache 量化功能,支持 int2 至 int4 精度,可显著降低长上下文生成的显存占用。实验显示 int4 精度在保持模型质量的同时,相比 fp16 可实现约 2.5 倍的显存节省,在 80GB A100 上支持最高 128k tokens 上下文。

推荐理由

官方在 Transformers 中集成 KV Cache 量化功能,提供 int2 至 int4 精度选项,显著降低长文本生成的显存占用。

来源:Hugging Face Blog · huggingface.co