YuE2-3B: open song generation that exposes the score as an interface
yue2
The open song generation model from m-a-p, 3B parameters, weights under CC-BY-NC-4.0, turning lyrics and a style prompt into a complete song with vocals and accompaniment at 48 kHz stereo. We list it as the open-weight top row for music in the audio domain on the strength of a combination that is close to unique among its peers: open weights plus an editable intermediate representation. One AR-NAR Mixture-of-Transformers backbone writes an ABC score (melody and chords) and semantic tokens, flow matching then produces acoustic latents and a VAE decodes them, so the score is an artefact that a person or an agent can read and edit instead of a black box whose only control is another sample. Three cot modes (full, melody, off) map onto composing, covering and direct generation, and the official agentic editing demo runs nine turns across fourteen versions from Mandarin pop to English jazz. Read the benchmark protocol carefully: on 192 WildSongBench prompts the best-of-8 SongBench average of 6.9632 sits above Suno v5 at 6.8721 and Suno v6 at 6.5562, but that is eight candidates with selection against a delivered single candidate; MuLan and AllMusicCaps, the two style-text alignment measures, still favour Suno v5, and PER at 8.44 percent trails Suno v6 Wild at 7.45 and MiniMax Music 3 at 6.27. Cost is the most concrete advantage: a 3.6 minute song in 71 seconds on an RTX 4090 24GB with a peak of 11.18 GiB, and 373 songs per hour on an H800 with vLLM at AR concurrency 32. Limits: non-commercial licence; identity preservation in covers comes almost entirely from a supplied score (CLEWS mAP collapses to 0.006 without one); the benchmark runs on the legacy VAE while the default release is the newer one; no technical report yet and no arena result. Graded C (vendor claim), not recomputed by us.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- WildSongBench SongBench 均分(best-of-8)
- Vendor Claim · 2026-09
- MATURITY
- Research
- research → demo → product → production
Our takeWe grade it C (vendor claim). WildSongBench publishes the benchmark itself, the 192 prompts, protocol notes for all nine metrics and a full precision CSV, which is far more transparent than a typical vendor post and makes this entry more trustworthy than others at the same grade. But the benchmark is operated by the model authors and the numbers are submitted by the model authors. We have not recomputed a single value and have not run a blind listening test, so contract section 13.5 forbids an upgrade to A. The fact that best-of-8 and delivered single candidates from Suno are different protocols is stated in their own notes, and we pass that through rather than smoothing it away.
What earns the top of the audio ladder is not the number one score but a combination that is close to unique in music generation: open weights plus an editable intermediate representation. Suno, Eleven Music and MiniMax Music are services where the only controls are the prompt and the section boundaries, and changing a chord progression means drawing another sample. YuE2 exposes the ABC score as an artefact, so reharmonisation, restyling a Mandarin pop song into English jazz, or making a solo quote an existing melody become versioned edits instead of a reroll. The nine step, fourteen version agentic editing chain is the strongest evidence for that route, and it matters beyond music: an agent that can edit a score means song generation plugs into the same read the artefact, edit the artefact, render again loop that this site catalogues under agent harness. Add a 3.6 minute song in 71 seconds on one 24GB card and the self-hosting cost lands at workstation scale, which no closed service offers.
Three hard limits govern selection. The licence is CC-BY-NC-4.0, non-commercial, so any commercial product either negotiates rights or switches service; a high score does not route around that. Covers depend on an external score: with no score supplied, the identity metric collapses to 0.006 mAP, so the real prerequisite for an AI cover pipeline is a working transcription chain (SheetSage2 plus ASR for the lyrics), not the model. Lyric intelligibility is still the weak spot, with PER at 8.44 percent behind Suno v6 Wild at 7.45 and MiniMax Music 3 at 6.27, which means lyric-dense passages need an audition before production. Note also the measurement protocol: 6.9632 comes from YuE2-Vae-legacy while the default release is YuE2-Vae, which sounds better and scores lower, so a reproduction run and a local listen are not the same decoder. The absent technical report and unverifiable training data, together with an arena that has published no result, are all folded into the C grade.
The problem it solves: making the intermediate state of music generation an editable score
YuE2-3B is the open song generation model from multimodal-art-projection (m-a-p). Give it lyrics and a style prompt and it returns a complete song with vocals and accompaniment at 48 kHz stereo, with weights released under CC-BY-NC-4.0. What separates it from the other music models in this catalogue is not audio quality but the intermediate representation: one AR-NAR Mixture-of-Transformers backbone first writes an ABC score (melody plus chords) and semantic tokens, flow matching then produces acoustic latents, and a VAE decodes them into stereo audio. The score therefore becomes an interface that a person or an agent can read and edit, rather than a black box whose only control is another sample.
That design maps directly onto three modes: cot="full" (melody and chord planning, the default), cot="melody" (melody-only planning, recommended for covers) and cot="off" (direct generation with no symbolic plan). You can also supply your own ABC, or call pipe.plan() to obtain the plan, edit it, then run generate_semantic(), synthesize() and decode() as separate stages.
WildSongBench: which rulers it wins and which it loses
The official comparison covers 192 WildSongBench prompts against 17 settings (2026-09-12 revision), with the SongBench average as the headline metric:
| Model | Open | Musicality ↑ | SongBench Avg ↑ | MuLan ↑ | AllMusicCaps ↑ | Q3O ↑ | PER ↓ |
|---|---|---|---|---|---|---|---|
| YuE2 (best-of-8) | Yes | 6.2666 | 6.9632 | 0.5051 | 0.3980 | 4.7009 | 9.79% |
| YuE2 (two candidates, lower PER) | Yes | 5.9075 | 6.7316 | 0.5068 | 0.4054 | 4.6819 | 8.44% |
| Suno v5 | No | 5.9918 | 6.8721 | 0.5428 | 0.4353 | 4.5907 | 8.10% |
| Suno v6 / v6 Wild | No | 5.6558 / 5.5644 | 6.5562 / 6.4195 | 0.4916 / 0.4999 | 0.4305 / 0.4316 | 4.6258 / 4.5898 | 7.58% / 7.45% |
| Mureka 9 | No | 6.0488 | 6.9377 | 0.4394 | 0.4102 | 4.6368 | 11.69% |
| LeVo 2 / ACE-Step 1.5 (open, same generation) | Yes | 5.4590 / 5.1588 | 6.3247 / 6.0118 | 0.3542 / 0.4372 | 0.2680 / 0.3869 | 3.9458 / 4.5809 | 26.12% / 7.46% |
Three things have to be read carefully. First, best-of-8 draws eight candidates and selects by Musicality, then Q3O, then PER; it is not single-shot quality, and comparing it against delivered single candidates from Suno is a different protocol. The authors state this in the protocol notes, but it should not be read as beating Suno at equal compute. Second, it does not win every ruler. On MuLan and AllMusicCaps, the two measures of audio-to-style-text alignment, Suno v5 is clearly higher (0.5428 and 0.4353 against 0.5051 and 0.3980). The YuE2 advantage concentrates in Musicality, the SongBench composite and Q3O (prompt adherence). Third, PER at 8.44 to 9.79 percent is not the lowest in this field: Suno v6 Wild sits at 7.45 percent and MiniMax Music 3 at 6.27 percent, so lyric intelligibility remains the weak spot of open models.
The cover generation set (SHS100K, 948 works times two styles times two seeds, 3792 songs per method) shows what the symbolic plan is actually doing. With a full score YuE2 reaches CLEWS mAP 0.647 and Hit@1 71.3 percent; removing chords drops it to 0.598 and 67.3 percent; with no score at all it collapses to 0.006 and 0.3 percent, while ACE-Step 1.5 manages only 0.024. Preserving the identity of the source song is therefore carried almost entirely by an externally supplied score (transcribed with SheetSage2), not by any implicit audio memory inside the model. That is both the capability and the limit: you need the score to cover the song.
Speed and memory: one 24GB card, which is the real difference from a closed service
| Setup | Mode | LM tokens/s | Generation / audio | Peak VRAM |
|---|---|---|---|---|
| RTX 4090 24GB (HF package, PyTorch + CUDA graphs + FlashAttention) | full | 139.48 | 71.04s / 214.85s for a 3.6 minute song | 11.18 GiB |
| RTX 4090 24GB | off | 121.07 | 57.91s / 196.88s | 11.09 GiB |
| H800 80GB + vLLM 0.19, AR concurrency 16 | full | 2418.63 | 340.38 songs/hour | 78.66 GiB |
| H800 80GB + vLLM 0.19, AR concurrency 32 | full | 3231.74 | 373.53 songs/hour | 76.61 GiB |
The requirements are stated plainly: Linux, Python 3.10+, a 24GB NVIDIA GPU with BF16 support, 24GB of free host RAM, no quantisation, one song at a time. Maximum-context testing peaked at 14.08 GiB. The serving row is a separate vLLM runtime, and 373 songs per hour is warm batch throughput rather than request latency, so the two figures should not be mixed.
Agentic editing: the official demo is a nine step chain across fourteen versions
Editing a song through an agent is treated as a first class workflow. The Last Train moves from Mandarin pop to English jazz with modern harmony and a saxophone solo built around two complete statements of Twinkle Twinkle Little Star, across nine turns and fourteen versions, with the conversation, score, prompt, lyrics and full audio published at every step. The engineering implication is that the agent receives three readable and editable artefacts, score.abc plus the original prompt plus the lyrics, and hands them back to YuE2 to render, rather than pressing generate again. For strict reharmonisation the authors require the agent to preserve melody pitches and rhythm and to check every sustained note against the new chords.
Limits: licence, VAE protocol and a technical report that is not out yet
- CC-BY-NC-4.0 is non-commercial. However good the benchmark, the weights cannot go straight into a commercial product. Commercial use needs a separate licence, or a paid service such as Suno or Eleven Music where the subscription carries the rights. This is the hardest line between YuE2 and every closed entry in this domain.
- The benchmark runs on YuE2-Vae-legacy while the default release is YuE2-Vae. The authors state that the legacy decoder scores higher on musicality while the newer one sounds better. The 6.9632 in that table is therefore not the decoder you get from a default local install, and reproducing it requires passing
vae="m-a-p/YuE2-Vae-legacy"explicitly. - No technical report yet. The citation points at the earlier YuE paper (arXiv 2503.08638), so training data composition, data scale and post-training details cannot be verified at present.
- The blind arena is still collecting votes. Music Arena runs anonymous preference voting against leading proprietary models and has no published result. Automatic metrics and human preference diverge more in music than in most modalities, so until those votes land, beating Suno remains a benchmark statement rather than a listening one.