PaperCompiler:通过仓库级规格编译实现忠实的论文到代码生成

内容摘要
概述: 论文到代码的生成一直是一个挑战,因为论文通常以高层次描述方法,隐含实现假设,并要求生成的代码库保持方法逻辑、评估协议和跨文件一致性。为了解决这些问题,研究人员提出了PaperCompiler,一个将论文中的证据编译成明确的代码库级别实现规范的框架。 要点: 1. PaperCompiler通过编译论文中的证据,生成明确的代码库级别实现规范,以解决论文到代码生成中的挑战。 2. 该框架在保留源证明的同时,区分了论文支持的、推断的、外部委托的和未解决的信息。 3. 生成的规范包括非降级要求、所有权分配、跨文件依赖和文件级约束。 4. 代码库生成过程遵循这些编译后的规范,同时保持对论文未固定的本地工程选择的灵活性。 5. 在Paper2CodeBench基准测试中,PaperCompiler优于强基线,实现了13.8%的相对改进,在基于参考的忠实度方面从3.64提升到4.15,并减少了高严重性评估者的批评(从13.2%降至6.1%)。
概述:
论文到代码的生成一直是一个挑战,因为论文通常以高层次描述方法,隐含实现假设,并要求生成的代码库保持方法逻辑、评估协议和跨文件一致性。为了解决这些问题,研究人员提出了PaperCompiler,一个将论文中的证据编译成明确的代码库级别实现规范的框架。

要点:
1. PaperCompiler通过编译论文中的证据,生成明确的代码库级别实现规范,以解决论文到代码生成中的挑战。
2. 该框架在保留源证明的同时,区分了论文支持的、推断的、外部委托的和未解决的信息。
3. 生成的规范包括非降级要求、所有权分配、跨文件依赖和文件级约束。
4. 代码库生成过程遵循这些编译后的规范,同时保持对论文未固定的本地工程选择的灵活性。
5. 在Paper2CodeBench基准测试中,PaperCompiler优于强基线,实现了13.8%的相对改进,在基于参考的忠实度方面从3.64提升到4.15,并减少了高严重性评估者的批评(从13.2%降至6.1%)。

Faithfully translating research papers into repository-level implementations remains challenging because papers often describe methods at a high level, leave implementation assumptions implicit, and require generated repositories to preserve method logic, evaluation protocols, and cross-file consistency. Despite recent advances in paper-to-code agents, their intermediate outputs are often presented as free-form plans or summaries that downstream coding agents may ignore, reinterpret, or compress, leading to algorithmic simplification and inconsistent repository structure. To address these challenges, we introduce PaperCompiler, a paper-to-code generation framework that compiles paper-grounded evidence into explicit repository-level implementation specifications. PaperCompiler grounds implementation-relevant evidence while preserving source provenance and distinguishing paper-supported, inferred, externally delegated, and unresolved information. The resulting specifications encode non-degradation requirements, ownership assignments, cross-file dependencies, and file-level constraints. Repository generation proceeds under these compiled specifications while retaining flexibility over local engineering choices not fixed by the paper. PaperCompiler outperforms strong baselines on Paper2CodeBench, achieving a 13.8% relative improvement in reference-based fidelity (from 3.64 to 4.15) and reducing high-severity evaluator critiques (from 13.2% to 6.1%).

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值