跳到正文
原文
Hugging Face Blog·· 2022-01-13AI 评分13

Hugging Face Infinity 在 Intel Xeon CPU 上实现毫秒级延迟案例研究

Case Study: Millisecond Latency using Hugging Face Infinity and modern CPUs

AI 导读

Hugging Face Infinity 容器化方案在 Amazon EC2 C6i(Intel Ice Lake)实例上运行 DistilBERT,相比 vanilla transformers 实现最高 800% 的延迟与吞吐量提升。该方案支持特征提取、排序等任务,通过端到端优化管道降低 Transformer 模型部署成本。

来源:Hugging Face Blog · huggingface.co