Environment Evolution:为终端 Agent 提供持续学习信号的环境进化方法

内容摘要
本文提出了一种名为“环境进化”的方法,旨在为终端 Agent 提供持续学习信号。该方法通过增量提高环境难度,并在训练过程中逐步生成进化的环境,以提供连续的学习信号。具体而言,文章从多轮学习目标中推导出三个影响环境难度的进化方向,并通过一个循环式多智能体 harness 实现这些进化方向。实验结果表明,环境进化能够持续产生更困难的环境,并在 Qwen3.6-27B 和 Qwen3.6-35B-A3B 上通过简单的长时程强化学习训练验证了其有效性,分别提高了 14.4 和 18.0 个百分点的性能。
本文提出了一种名为“环境进化”的方法,旨在为终端 Agent 提供持续学习信号。该方法通过增量提高环境难度,并在训练过程中逐步生成进化的环境,以提供连续的学习信号。具体而言,文章从多轮学习目标中推导出三个影响环境难度的进化方向,并通过一个循环式多智能体 harness 实现这些进化方向。实验结果表明,环境进化能够持续产生更困难的环境,并在 Qwen3.6-27B 和 Qwen3.6-35B-A3B 上通过简单的长时程强化学习训练验证了其有效性,分别提高了 14.4 和 18.0 个百分点的性能。

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-03 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值