Hugging Face Blog·· 2022-08-17精选AI 评分68
Hugging Face 将 LLM.int8() 集成至 transformers 库
A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
AI 导读
Hugging Face 将 LLM.int8() 8-bit 量化技术集成至 transformers 和 accelerate 库,使 BLOOM-176B 等超大模型显存占用减半且性能无退化。该方案通过分离异常值计算,在 Turing 和 Ampere 架构 GPU 上实现高效推理。
推荐理由
官方将LLM.int8()集成至transformers库,提供降低显存占用且保持性能的方法与代码示例。
来源:Hugging Face Blog · huggingface.co