Hugging Face Blog·· 2024-03-15AI 评分38
Hugging Face 展示 Optimum Intel 与 fastRAG 在 CPU 上优化 BGE 嵌入模型
CPU Optimized Embeddings with 🤗 Optimum Intel and fastRAG
AI 导读
Hugging Face 演示如何利用 Optimum Intel 和 fastRAG 在 Xeon CPU 上加速 BGE 嵌入模型。通过量化等技术,该方案显著降低延迟并提升吞吐量,适用于 RAG 管道中的检索与重排序。BGE small、base 和 large 模型分别拥有 45M、110M 和 355M 参数,编码向量维度为 384/768/1024。
来源:Hugging Face Blog · huggingface.co