跳到正文
原文
Hugging Face Blog·· 2024-01-30AI 评分40

StarCoder 在 Intel Xeon 上结合 Q8/Q4 量化与推测解码实现加速

Accelerate StarCoder with 🤗 Optimum Intel on Xeon: Q8/Q4 and Speculative Decoding

AI 导读

StarCoder-15B 模型在 Intel 4th Gen Xeon 上结合 8bit 和 4bit 量化与推测解码,实现超过 7 倍的推理加速。Q8-StarCoder 在 HumanEval 上无精度损失,TTFT 和 TPOT 分别提速约 2.19 倍和 2.20 倍。INT4 量化通过减少模型加载时间进一步缓解内存带宽瓶颈。

来源:Hugging Face Blog · huggingface.co