YuE2 发布:融合象征性规划的音乐生成模型,SongBench 得分 6.9632 超越 Suno v5

内容摘要
概述: YuE2是一款音乐生成模型,它在SongBench测试中取得了6.9632的高分,超越了Suno v5的6.8721。该模型通过将符号化的乐谱转换为音乐标记、声学潜力和完整歌曲音频,支持音乐创作、覆盖和编辑。YuE2的架构包括约3.59亿个参数和28层,同时MERT2模型在音乐表征方面也取得了显著成果,在MARBLE测试中达到14项指标中的SOTA水平。 要点: 1. YuE2在SongBench测试中得分6.9632,超越Suno v5的6.8721。 2. YuE2模型支持音乐创作、覆盖和编辑,具有约3.59亿个参数和28层。 3. MERT2模型在音乐表征方面取得显著成果,在MARBLE测试中达到14项指标中的SOTA水平。 4. SheetSage2模型将录音转换为可编辑的乐谱,并在多个音乐转录任务中达到SOTA水平。 5. YuE2和MERT2模型主要在CC0音乐和合成数据上训练,Tokenwave.AI提供了大部分合成训练数据。
概述:
YuE2是一款音乐生成模型,它在SongBench测试中取得了6.9632的高分,超越了Suno v5的6.8721。该模型通过将符号化的乐谱转换为音乐标记、声学潜力和完整歌曲音频,支持音乐创作、覆盖和编辑。YuE2的架构包括约3.59亿个参数和28层,同时MERT2模型在音乐表征方面也取得了显著成果,在MARBLE测试中达到14项指标中的SOTA水平。

要点:
1. YuE2在SongBench测试中得分6.9632,超越Suno v5的6.8721。
2. YuE2模型支持音乐创作、覆盖和编辑,具有约3.59亿个参数和28层。
3. MERT2模型在音乐表征方面取得显著成果,在MARBLE测试中达到14项指标中的SOTA水平。
4. SheetSage2模型将录音转换为可编辑的乐谱,并在多个音乐转录任务中达到SOTA水平。
5. YuE2和MERT2模型主要在CC0音乐和合成数据上训练,Tokenwave.AI提供了大部分合成训练数据。

From score to song

Listen to a song, then explore the melody, rhythm, and chords in its symbolic plan.

Selected score

Loading selected song…

The generated song

The symbolic score

Original score recording

Interactive ABC score

Red notes follow the score recording.

View original score pages

All selected scores

Cover & Editing

A familiar song can take a different shape. Listen to changes in melody, lyrics, tempo, and arrangement.

Agentic music editing

A song takes shape through a conversation. Loading the editing story…

Genre Explorer

The listening selection, gathered across genres and languages.

Model & Results

YuE2 (best-of-8) reaches 6.9632 on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench (192 prompts). Suno v5 scores 6.8721 in the same comparison.

Figure 1

YuE2 at the frontier.

Vector PDF
Paper Figure 1: a WildSongBench comparison of song quality and text alignment. YuE2 and YuE2 best-of-8 are competitive with the evaluated proprietary systems. Bubble area represents AudioBox production quality.
Song quality and text alignment on WildSongBench. Bubble area shows AudioBox PQ; black outlines mark Pareto optima on the two plotted axes. Bo8 = best-of-8. How to read the indices.

Model architecture

Composing in symbols, performing in audio.

Vector PDF
YuE2 architecture: an AR–NAR Mixture-of-Transformers turns an editable symbolic score into semantic tokens, acoustic latents, and full-song audio. The same generator supports creation, covering, and editing.
An editable score becomes semantic music tokens, acoustic latents, and full-song audio. The model has approximately 3.59B parameters and 28 layers, and supports creation, covering, and editing. The AR and NAR experts share an attention computation while using separate normalization, projections, and MLPs.

Explore benchmark scores

WildSongBench · 15 settings · 7 metrics

WildSongBench

192 prompts

WildSongBench · SongBench
RankSystem / settingPromptsSongBench ↑

September 5, 2026 evaluation · rankings vary by metric.

Download results

How to read these results

WildSongBench. 192 prompts and 15 system settings. The table reports automatic evaluation scores. Best-of-8 selects one of eight generations by musicality, prompt control, and lyric accuracy.

Figure 1. Song quality combines SongBench and SongEval; text alignment combines MuLan, AllMusicCaps, and prompt control. Both axes show normalized comparison indices. Bubble area represents AudioBox production quality.

MERT2 · Music representations

Learning the structure behind the sound.

Vector PDF
MERT2 architecture: offline target synthesis combines MuQ and Qwen2-Audio features into four code streams. A ConvNeXt frontend and 24-layer Conformer learn by masked prediction, then branch into full-song representations for SheetSage2 and a causal tokenizer curriculum for YuE2.
MERT2 provides the music representations behind both analysis and generation. A ConvNeXt frontend and 24-layer Conformer learn to predict masked codes built from complementary MuQ and Qwen2-Audio features. Full-song adaptation supplies SheetSage2 with musical context; a separate causal branch becomes YuE2’s 25-Hz semantic tokenizer.

State of the art on MARBLE

MERT2-30s and MERT2-FS (full-song) achieve SOTA on 14 of 15 MARBLE metrics, leading across tagging, key, genre, and emotion recognition.

SOTA metricsMERT2-30s & MERT2-FS
14 / 15
Genre accuracy · GTZANMERT2-30s · score × 100
91.72
Key refined accuracy · GiantStepsMERT2-FS · score × 100
67.05

Explore MERT2 benchmark scores

MARBLE · 15 metrics · 2 encoders

Scores × 100 · higher is better · bold marks the highest displayed score.
Benchmark / metricBest published baselineMERT2-30sMERT2-FS
MTT · TaggingROC-AUC91.70PupuJEPA-Large91.9191.74
MTT · TaggingAverage precision40.80PupuJEPA-Large41.2941.20
GiantSteps · KeyRefined accuracy66.10PupuJEPA-Large66.9767.05
GTZAN · GenreAccuracy86.90PupuJEPA-Large91.7290.69
GTZAN · BeatF191.00PupuJEPA-Large90.5990.57
EmoMusic · Valence62.50PupuJEPA-Large63.2363.52
EmoMusic · Arousal78.50PupuJEPA-Huge80.0178.14
MTG-Jamendo · InstrumentROC-AUC78.40PupuJEPA-Large80.2780.27
MTG-Jamendo · InstrumentAverage precision21.20PupuJEPA-Large22.8923.51
MTG-Jamendo · Mood / themeROC-AUC76.20PupuJEPA-Large79.4478.74
MTG-Jamendo · Mood / themeAverage precision15.50Dasheng-1.2B16.6815.74
MTG-Jamendo · GenreROC-AUC86.30AudioMAE++88.0187.98
MTG-Jamendo · GenreAverage precision20.10PupuJEPA-Large / PupuJEPA-Huge21.2220.66
MTG-Jamendo · Top 50ROC-AUC83.10AudioMAE++ / PupuJEPA-Huge84.1884.13
MTG-Jamendo · Top 50Average precision31.10AudioMAE++32.1731.62

SOTA counts use the best score across the two MERT2 encoders against the nine published baselines in this comparison. Both encoders have 632M parameters. MERT2-30s uses a 30-second training context; MERT2-FS uses 300 seconds. These are full-context representation benchmarks. MERT2 reports the best observed results across representations selected using test scores; each ROC-AUC / AP pair uses the same representation.

Published baseline comparison ↗

Download all 11 models

SheetSage2 · Audio to score

Hear a song. Read its composition.

Vector PDF
SheetSage2 architecture: a full-song MERT2-FS encoder with trainable adapters feeds a six-layer autoregressive decoder. Task prompts select beat, section, key, chord, and melody events, which share a timeline and become ABC notation and a lead sheet.
SheetSage2 turns a recording into an editable lead sheet. A full-song MERT2-FS encoder, adapted with LoRA, feeds a six-layer autoregressive decoder. Task prompts select beats, sections, keys, chords, and melodies; a shared event timeline becomes ABC notation with vocal and instrumental melody voices. These scores supply symbolic training targets for YuE2.

Six transcription tasks, one model

SheetSage2 achieves SOTA on 10 of 13 benchmark metrics with one model for beat, downbeat, key, chord, structure, and melody transcription.

SOTA metricsOne model · six transcription tasks
10 / 13
Vocal melody · RWC-PopPitch-class note F1 · score × 100
82.51
Chord recognition · osu2017Maj/min · score × 100
90.08

Explore SheetSage2 benchmark scores

6 tasks · 13 metrics

Scores × 100 · higher is better · bold marks the highest displayed score.
Benchmark / taskMetricBest comparisonSheetSage2
GTZANBeatF188.75Beat This!85.65
osu2017BeatF191.55Madmom92.29
GTZANDownbeatF178.28Beat This!79.51
osu2017DownbeatF184.99Beat This!91.97
GiantStepsKeyWeighted score74.62Madmom77.73
GTZANKeyWeighted score74.43S-KEY75.77
osu2017ChordMaj/min84.59Jiang et al. 201990.08
Chords1217ChordMaj/min84.09ChordFormer83.81
HarmonixSetStructureAccuracy80.03SongFormer80.51
HarmonixSetStructureBoundary F1 · 0.5 s70.63SongFormer67.96
HarmonixSetStructureBoundary F1 · 3 s79.50SongFormer82.86
RWC-PopMelodyVocal pitch-class F162.71SheetSage182.51
RWC-PopMelodyFull pitch-class F164.02SheetSage175.29

SOTA counts refer to the leading scores against SheetSage1, Madmom, and the task-specific systems in this comparison. Results use one model selected by validation loss. Melody F1 measures pitch-class notes; structure F1 measures section boundaries at the stated tolerance. On Chords1217, ChordFormer uses five-fold cross-validation, while SheetSage2 evaluates one fixed model on all 1,217 tracks.

Full comparison includes SheetSage1, Madmom, and task-specific systems.

Download all results

Training data

Our models are trained primarily on CC0 music and synthetic data. Tokenwave.AI provides most of our synthetic training data under license. We are committed to the ethical and responsible use of data.

MERT2
700K hours
SheetSage2
28.4K hours
YuE2
346K hours

原始发布方:Hacker News 热门(buzzing.cc 中文翻译)

原文时间:2026-09-11 22:20:57 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值