跳到正文
原文
Hugging Face Blog·· 2025-05-13精选AI 评分62

Hugging Face Inference Endpoints 上线极速 Whisper 部署

Blazingly fast whisper transcriptions with Inference Endpoints

AI 导读

Hugging Face 在 Inference Endpoints 推出基于 vLLM 的 Whisper 部署选项,利用 torch.compile、CUDA graphs 和 float8 KV cache 等技术实现近 8 倍推理加速。在保持转录准确率不变的前提下,显著提升了长音频处理效率,并提供 Python 代码与 FastRTC 实时演示供社区使用。

推荐理由

基于vLLM与多项底层优化实现近8倍推理加速,为语音转写提供低成本部署方案。

来源:Hugging Face Blog · huggingface.co