DisCo 提出 Repo-To-Skill 方法,将 1,000 个 ML 仓库蒸馏成 5,000+ 技能并提升智能体研究表现

内容摘要
概述: DisCo是一种基于技能的智能体研究工具,旨在将机器学习(ML)领域的知识转化为可重用的技能,以提升智能体在研究中的表现。该方法通过将广泛使用的ML仓库蒸馏成可验证的技能,从而实现知识的复用,避免每次运行时重新发现知识。 要点: 1. DisCo智能体通过创建和利用技能来提升研究表现。 2. 该方法分为两种形式:任务无关的蒸馏和任务导向的蒸馏。 3. AREX-Skill Library包含从1,000个ML仓库中蒸馏出的5,000多个验证技能,涵盖20个领域和178个能力家族。 4. 在固定GPT-5.5骨干、研究工具和下游执行预算的条件下,配备技能的智能体在MLE-bench、PaperBench、FrontierCS和PassNet上的表现分别提升了134.3%、34.4%、9.2%和14.0%。 5. 这些提升得益于在固定设置下添加了蒸馏的操作上下文。
概述:
DisCo是一种基于技能的智能体研究工具,旨在将机器学习(ML)领域的知识转化为可重用的技能,以提升智能体在研究中的表现。该方法通过将广泛使用的ML仓库蒸馏成可验证的技能,从而实现知识的复用,避免每次运行时重新发现知识。

要点:
1. DisCo智能体通过创建和利用技能来提升研究表现。
2. 该方法分为两种形式:任务无关的蒸馏和任务导向的蒸馏。
3. AREX-Skill Library包含从1,000个ML仓库中蒸馏出的5,000多个验证技能,涵盖20个领域和178个能力家族。
4. 在固定GPT-5.5骨干、研究工具和下游执行预算的条件下,配备技能的智能体在MLE-bench、PaperBench、FrontierCS和PassNet上的表现分别提升了134.3%、34.4%、9.2%和14.0%。
5. 这些提升得益于在固定设置下添加了蒸馏的操作上下文。

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值