跳到正文
原文
Hugging Face Blog·· 10 天前精选AI 评分80

Transformers 支持运行 llama.cpp GGUF 量化模型

Transformers now runs llama.cpp quants

AI 导读

Hugging Face Transformers 新增对 GGUF 格式模型的支持,允许用户通过熟悉的 API 在本地加载和运行 llama.cpp 的量化检查点。该功能目前主要面向 Apple Silicon 设备,通过复用 ggml 内核实现了接近 llama.cpp 的推理性能,并支持使用 transformers serve 暴露 OpenAI 兼容接口。

推荐理由

原文展示了在 transformers 中直接运行 GGUF 模型的方法及性能基准,为本地开发者提供了无需切换推理引擎即可复用现有 PyTorch 工作流的可行路径。

来源:Hugging Face Blog · huggingface.co