DiagEvo:基于分层错误记忆的诊断引导式自我进化

内容摘要
DiagEvo是一种基于分层错误记忆的诊断引导式自我进化的语言模型。该方法通过分析求解器的失败历史,提取重复出现的错误原因,并存储在分层错误原因记忆中。该记忆将相关原因分组在技能节点下,并跟踪每个原因的状态,根据在目标问题上的自我一致性将其标记为“活跃”或“掌握”。挑战者利用这些状态和复发次数来平衡针对原因的生成与自由探索。双信心过滤仅保留中等难度的题目,当最常见求解器答案有明确的投票优势时。DiagEvo从自我玩耍过程中产生的信息中推导出其课程,无需外部任务资源。在默认的4B诊断器下,DiagEvo在所有九个基准测试中,平均准确率均优于所有基线,包括Qwen3-4B、Qwen3-8B和OctoThinker-8B。在Qwen3-8B上,它在五个数学推理基准测试中的平均准确率达到72.3%,比R-Zero高4.5个百分点。在所有九个基准测试中的平均准确率为57.4%,比DARC高1.1个百分点。消融实验表明,分层错误原因记忆和双信心过滤都对这些收益有贡献。
DiagEvo是一种基于分层错误记忆的诊断引导式自我进化的语言模型。该方法通过分析求解器的失败历史,提取重复出现的错误原因,并存储在分层错误原因记忆中。该记忆将相关原因分组在技能节点下,并跟踪每个原因的状态,根据在目标问题上的自我一致性将其标记为“活跃”或“掌握”。挑战者利用这些状态和复发次数来平衡针对原因的生成与自由探索。双信心过滤仅保留中等难度的题目,当最常见求解器答案有明确的投票优势时。DiagEvo从自我玩耍过程中产生的信息中推导出其课程,无需外部任务资源。在默认的4B诊断器下,DiagEvo在所有九个基准测试中,平均准确率均优于所有基线,包括Qwen3-4B、Qwen3-8B和OctoThinker-8B。在Qwen3-8B上,它在五个数学推理基准测试中的平均准确率达到72.3%,比R-Zero高4.5个百分点。在所有九个基准测试中的平均准确率为57.4%,比DARC高1.1个百分点。消融实验表明,分层错误原因记忆和双信心过滤都对这些收益有贡献。

Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can plateau or decline across rounds. Unguided methods steer question generation with signals such as difficulty, learnability, or diversity. These signals keep questions challenging and varied but do not specify which unresolved reasoning weaknesses later rounds should target. Guided methods obtain direction from external task resources, including human examples, document corpora, or specified difficulty targets, and therefore rely on task information supplied outside the self-play loop. We show that the needed direction can instead be derived from the solver's own failure history. We introduce DiagEvo, whose diagnostician extracts recurring error causes from this history and stores them in a hierarchical error-cause memory. The memory groups related causes under skill nodes and tracks each as Active or Mastered according to self-consistency on targeted questions. The challenger uses these states and recurrence counts to balance cause-targeted generation with free exploration. Double-confidence filtering retains intermediate-difficulty questions only when the most common solver answer has a clear vote lead. DiagEvo derives its curriculum from information produced during self-play, without external task resources. With the default 4B diagnostician, DiagEvo outperforms every baseline in mean accuracy across all nine benchmarks for each of the three solvers: Qwen3-4B, Qwen3-8B, and OctoThinker-8B. On Qwen3-8B, it reaches 72.3% mean accuracy across five mathematical reasoning benchmarks, 4.5 percentage points above R-Zero. Its mean accuracy across all nine benchmarks is 57.4%, 1.1 percentage points above DARC. Ablations show that the hierarchical error-cause memory and double-confidence filtering both contribute to these gains.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值