Harness-of-Harness 论文提出多天自主软件开发框架,平均相对提升 52.25%

内容摘要
该论文研究了自主软件开发,提出了一种名为 Harness-of-Harness(HoH)的框架,该框架允许编码代理在自主开发过程中持续改进软件。HoH 在现有的编码代理 Harness 上运行,将它们的执行组织成迭代规划-编码-测试循环。为了在循环中维持改进,HoH 平衡修复与能力增长,将开发范围缩小到小而可验证的增量,将实现时测试与独立评估分离,并限制可验证输出而不是规定代理工作流程。它逐步暴露交付成果、特定角色的工具和技能,鼓励重用而不是重新创造,并维护版本化的项目历史。在 GameCraft-Bench、FrontierSWE 和 ProgramBench 上,与三个 Harness 模型对(Codex 与 GPT-5.5、OpenCode 与 DeepSeek-V4-Pro、Pi 与 MiniMax-M3)相比,HoH 在三个迭代后平均相对提升 52.25%,最大提升 82.86%。在超过 70 次迭代的为期数天的部署中,HoH 自主开发了一款第一人称射击游戏,具有连贯的故事情节、完全实现的核心机制、可玩性、精美的视觉效果和集成音频。
该论文研究了自主软件开发,提出了一种名为 Harness-of-Harness(HoH)的框架,该框架允许编码代理在自主开发过程中持续改进软件。HoH 在现有的编码代理 Harness 上运行,将它们的执行组织成迭代规划-编码-测试循环。为了在循环中维持改进,HoH 平衡修复与能力增长,将开发范围缩小到小而可验证的增量,将实现时测试与独立评估分离,并限制可验证输出而不是规定代理工作流程。它逐步暴露交付成果、特定角色的工具和技能,鼓励重用而不是重新创造,并维护版本化的项目历史。在 GameCraft-Bench、FrontierSWE 和 ProgramBench 上,与三个 Harness 模型对(Codex 与 GPT-5.5、OpenCode 与 DeepSeek-V4-Pro、Pi 与 MiniMax-M3)相比,HoH 在三个迭代后平均相对提升 52.25%,最大提升 82.86%。在超过 70 次迭代的为期数天的部署中,HoH 自主开发了一款第一人称射击游戏,具有连贯的故事情节、完全实现的核心机制、可玩性、精美的视觉效果和集成音频。

This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development. HoH operates on existing coding-agent harnesses, and organizes their executions into iterative planning-coding-testing loops. To sustain improvement across loops, HoH balances repair with capability growth, scopes development into small and verifiable increments, separates implementation-time testing from independent evaluation, and constrains verifiable outputs rather than prescribing agent workflows. It progressively exposes deliverables, role-specific tools, and skills, encourages reuse rather than recreation, and maintains versioned project histories. On GameCraft-Bench, FrontierSWE, and ProgramBench, three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), HoH consistently outperforms the corresponding standalone harnesses, achieving an average relative gain of 52.25 percent and a maximum gain of 82.86 percent after three iterations. In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio. Github: https://github.com/Flesymeb/HarnessOfHarness Project Page: https://flesymeb.github.io/HarnessOfHarness/

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值