Hugging Face Blog·· 2024-05-09AI 评分23
使用 Intel Gaudi 2 和 Xeon 构建成本高效的 RAG 应用
Building Cost-Efficient Enterprise RAG applications with Intel Gaudi 2 and Intel Xeon
AI 导读
Intel 结合 Gaudi 2 加速器和 Xeon CPU,通过 OPEA 平台与 LangChain 框架,展示了构建企业级 RAG 应用的方法。其中嵌入模型运行于 Granite Rapids CPU,LLM 运行于 Gaudi 2,并支持通过 TGI 部署 Llama 等模型。测试显示,启用 FP8 量化相比 BF16 可获得 1.8 倍的吞吐量提升。
来源:Hugging Face Blog · huggingface.co