Nemotron-3 系列后训练模型在 IOI 编程竞赛中超过人类最高分

内容摘要
Nemotron-3 系列后训练模型在 IOI 编程竞赛中取得了显著成绩。通过结合大规模问题编纂、合成推理轨迹、监督微调和强化学习等手段,Nemotron-3-Nano-CC 和 Nemotron-3-Ultra-CC 在 IOI 2025 中分别取得了 291 分和 502 分,超越了人类最高分。此外,引入的 GenCorrect 策略进一步提升了模型的性能。在 IOI 2026 中,Nemotron-3-Ultra-CC 系统在相同条件下取得了 535.4 分,超过了人类最高分 498.27 分,成为首个在 IOI 问题集中超越人类最高分的人工智能系统。
Nemotron-3 系列后训练模型在 IOI 编程竞赛中取得了显著成绩。通过结合大规模问题编纂、合成推理轨迹、监督微调和强化学习等手段,Nemotron-3-Nano-CC 和 Nemotron-3-Ultra-CC 在 IOI 2025 中分别取得了 291 分和 502 分,超越了人类最高分。此外,引入的 GenCorrect 策略进一步提升了模型的性能。在 IOI 2026 中,Nemotron-3-Ultra-CC 系统在相同条件下取得了 535.4 分,超过了人类最高分 498.27 分,成为首个在 IOI 问题集中超越人类最高分的人工智能系统。

Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-02 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值