Hugging Face Blog·· 2024-08-14AI 评分37
Hugging Face 复现 Infini-Attention 失败:压缩越多性能越差,Ring Attention 仍是最佳方案
A failed experiment: Infini-Attention, and why we should keep trying?
AI 导读
Hugging Face 复现 Infini-Attention 失败,发现随着内存压缩次数增加,模型性能反而下降。在将 Llama 3 8B 扩展至 100 万 token 上下文的实验中,Ring Attention、YaRN 和 rope scaling 仍是目前延长预训练模型上下文长度的最佳方案。
来源:Hugging Face Blog · huggingface.co