推理预填充实验显示 Qwen3.8 与 GPT-5.5 Pro 重合度提升 18.18 个百分点

内容摘要
实验结果显示,在推理预填充实验中,Qwen3.8模型在GPT-5.5 Pro的指导下,其与GPT-5.5 Pro的重合度提升了18.18个百分点。实验中,针对每个问题,从目标模型中生成两种回答:一种是普通未预填充的回答,另一种是在目标模型的推理通道中插入GPT-5.5 Pro推理的前1%的回答。结果显示,Qwen3.8模型在STEM、非STEM和合成谜题类别中均表现出显著提升。此外,Kimi K3模型在有无预填充的情况下与GPT-5.5 Pro的重合度最高,分别为31.11%和35.65%。
实验结果显示,在推理预填充实验中,Qwen3.8模型在GPT-5.5 Pro的指导下,其与GPT-5.5 Pro的重合度提升了18.18个百分点。实验中,针对每个问题,从目标模型中生成两种回答:一种是普通未预填充的回答,另一种是在目标模型的推理通道中插入GPT-5.5 Pro推理的前1%的回答。结果显示,Qwen3.8模型在STEM、非STEM和合成谜题类别中均表现出显著提升。此外,Kimi K3模型在有无预填充的情况下与GPT-5.5 Pro的重合度最高,分别为31.11%和35.65%。

Reasoning prefills on a few open models, v1.1

A follow-up to Reasoning prefills on a few open models and Stolen Thoughts

This v1.1 reruns the reasoning-prefill experiment with GPT-5.5 Pro as the teacher.

For each problem, I generated two responses from each target model:

  1. an ordinary, unprefilled response; and
  2. a response starting with the first 1% of GPT-5.5 Pro's reasoning, inserted into the target model's reasoning channel.

The visible answer remained freely generated. I then measured how much of the teacher's visible answer appeared in the first 100 tokens of the target model's answer. As in the previous post, each score is the mean of unigram, bigram, and trigram source recall. Deltas are absolute percentage-point changes.

All problems

The evaluation contains 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles.

Model n Unprefilled GPT-5.5 Pro reasoning prefill Delta
DeepSeek V4 Flash 45 27.30% 26.13% −1.17 pp
Inkling 45 19.99% 20.45% +0.46 pp
Kimi K3 45 31.11% 35.65% +4.54 pp
Qwen3.8 A95B 45 16.79% 34.97% +18.18 pp

Qwen by category

Category n Unprefilled GPT-5.5 Pro reasoning prefill Delta
STEM 15 19.26% 46.24% +26.99 pp
Non-STEM 15 20.62% 33.42% +12.80 pp
Puzzle 15 10.49% 25.23% +14.75 pp
All 45 16.79% 34.97% +18.18 pp

Discussion

Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +18.18 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.

Kimi K3 has the highest overlap with GPT-5.5 Pro both without and with the prefill (31.11% and 35.65%), although the prefill adds only +4.54 points.

原始发布方:Hacker News 热门(buzzing.cc 中文翻译)

原文时间:2026-09-10 04:46:59 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值