Cliff:从第一个错误学习过程奖励的 RLVR 奖励塑形策略

内容摘要
Cliff是一种基于可验证奖励的强化学习(RLVR)奖励塑形策略,旨在提升大型语言模型(LLM)的推理性能。该策略利用现成的LLM作为教师,识别每个rollout中的第一个错误,将rollout自然分解为正确的前缀和错误的后缀。Cliff将此信号转换为token级别的优势,为正确的前缀分配正优势,随后提供负面反馈。实验表明,Cliff在12个不同场景中均能提升推理性能,比在线策略蒸馏提高15%,比标准GRPO提高7%,即使教师能力有限。此外,研究还分析了“ground truth”在Cliff中的作用,并探讨了其训练动态。这些结果确立了Cliff作为一种简单、通用且有效的改进RLVR的方法。
Cliff是一种基于可验证奖励的强化学习(RLVR)奖励塑形策略,旨在提升大型语言模型(LLM)的推理性能。该策略利用现成的LLM作为教师,识别每个rollout中的第一个错误,将rollout自然分解为正确的前缀和错误的后缀。Cliff将此信号转换为token级别的优势,为正确的前缀分配正优势,随后提供负面反馈。实验表明,Cliff在12个不同场景中均能提升推理性能,比在线策略蒸馏提高15%,比标准GRPO提高7%,即使教师能力有限。此外,研究还分析了“ground truth”在Cliff中的作用,并探讨了其训练动态。这些结果确立了Cliff作为一种简单、通用且有效的改进RLVR的方法。

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值