Hugging Face Blog·· 2025-01-23AI 评分50
NVIDIA 推出 KVPress 工具包以压缩 LLM 的 KV Cache
Mastering Long Contexts in LLMs with KVPress
AI 导读
NVIDIA 发布 Python 工具包 KVPress,通过集成 KnormPress、SnapKVPress 等先进算法压缩 KV Cache,解决长上下文内存瓶颈。该方案支持在 transformers 库中结合量化技术,显著降低如 Llama 3-70B 处理 1M tokens 时高达 327.6GB 的显存需求,助力高效部署大语言模型。
来源:Hugging Face Blog · huggingface.co