Hugging Face Blog·· 2026-05-07AI 评分42
vLLM V1 在 PipelineRL 中实现与 V0 的 RL 训练对齐
vLLM V0 to V1: Correctness Before Corrections in RL
AI 导读
PipelineRL 通过修复 logprobs 语义、运行时默认值及 fp32 lm_head,使 vLLM V1 (0.18.1) 在 GSPO 训练中匹配 vLLM V0 (0.8.5) 参考基准。该方案消除了推理引擎与训练器间的 logprobs 偏差,确保 clip rate、KL 和 reward 等指标回归正常轨迹。
来源:Hugging Face Blog · huggingface.co