跳到正文
原文
Hugging Face Blog·· 2025-05-13AI 评分44

Hugging Face Inference Endpoints 推出极速 Whisper 部署方案

Blazingly fast whisper transcriptions with Inference Endpoints

AI 导读

Hugging Face 在 Inference Endpoints 上推出基于 vLLM 的 OpenAI Whisper 优化部署,性能提升高达 8x。该方案针对 NVIDIA L4/L40s GPU 进行 torch.compile、CUDA graphs 及 float8 KV cache 等底层优化,在保持转录质量的同时显著降低延迟。

来源:Hugging Face Blog · huggingface.co