跳到正文
原文
Hugging Face Blog·· 2023-11-07AI 评分30

使用 AWS Inferentia2 加速 Llama 2 生成

Make your llama generation time fly with AWS Inferentia2

AI 导读

Hugging Face 推出 optimum-neuron 工具,支持在 AWS Inferentia2 上部署 Llama 2 模型以加速文本生成。用户可通过 NeuronModelForCausalLM API 将 Llama-2-7b-hf 等模型编译导出至 Neuron 格式,并配置 num_cores、batch_size 及 sequence_length 参数优化推理性能。

来源:Hugging Face Blog · huggingface.co