SMELT:算力匹配下 MoE 循环 Transformer 的缩放律研究

内容摘要
Looped Transformers通过迭代共享的层块来增加有效深度,但大多数评估在固定模型大小下进行,将架构优势与额外的FLOPs混淆。该研究在MoE Transformers上研究了循环,同时紧密匹配每token的FLOPs、总非嵌入参数和KV缓存。通过一系列消融实验,研究者得出了一种名为SMELT(稀疏MoE Transformer,中间层循环两次)的方案,该方案在中间层的一半上循环两次,同时在所有三个预算上与未循环的基线匹配。SMELT在四个大小上进行了扩展,参数量高达54B,并为每个架构拟合了单独的Chinchilla-style缩放律。SMELT的损失随着计算能力的增加而更快地下降,在计算最优前沿上节省了6.8%至18.0%的训练FLOPs。这种优势超出了验证损失的预测,在代码上最大,并且随着样本长度和上下文示例数量的增加而增长。机制分析表明,第二次访问减少了注意力汇聚,并将质量重新导向与内容相关的token,这种归纳偏差可能是观察到的性能增益的基础。这些结果表明,即使在预算匹配的情况下,循环也可以提高Transformers,提供了一种将深度重用转化为可衡量收益的实用方案。
Looped Transformers通过迭代共享的层块来增加有效深度,但大多数评估在固定模型大小下进行,将架构优势与额外的FLOPs混淆。该研究在MoE Transformers上研究了循环,同时紧密匹配每token的FLOPs、总非嵌入参数和KV缓存。通过一系列消融实验,研究者得出了一种名为SMELT(稀疏MoE Transformer,中间层循环两次)的方案,该方案在中间层的一半上循环两次,同时在所有三个预算上与未循环的基线匹配。SMELT在四个大小上进行了扩展,参数量高达54B,并为每个架构拟合了单独的Chinchilla-style缩放律。SMELT的损失随着计算能力的增加而更快地下降,在计算最优前沿上节省了6.8%至18.0%的训练FLOPs。这种优势超出了验证损失的预测,在代码上最大,并且随着样本长度和上下文示例数量的增加而增长。机制分析表明,第二次访问减少了注意力汇聚,并将质量重新导向与内容相关的token,这种归纳偏差可能是观察到的性能增益的基础。这些结果表明,即使在预算匹配的情况下,循环也可以提高Transformers,提供了一种将深度重用转化为可衡量收益的实用方案。

Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值