Skip to content
← Tags

#Text-to-Music (2)

Speech & AudioMediaTopC

Suno: the platform that pulls the second half of music production inside one product

The most complete commercial platform on the text-to-song route: a style description, your own lyrics, a hummed melody or a recorded riff all work as a starting point, a full song with vocals and instrumentation arrives in seconds, and the same product then extends it, edits sections, restyles it, extracts stems and remasters it. We list it as the top closed entry for music in the audio domain on workflow completeness rather than peak audio quality. Most generative music products cover drafting and picking a take and stop there; Suno connects the rest through stem extraction with up to 12 stems, MIDI export and Suno Studio, and Studio speaks conventional DAW semantics rather than adding another prompt box, with a multitrack timeline, take lanes, comping, manual BPM to settle tempo drift and per clip transpose and speed. Custom Models trains up to three private style variants from six or more tracks you own, and Voices, formerly Personas, generates in your own singing timbre with a verification step. The tier split matters: the free plan covers creation (generation, lyrics, Cover, crop and fade, audio upload) while stems, Add Vocals, Voices and Custom Models require Pro or Premier and Studio is Premier only on desktop web, so real cost modelling should assume Premier. Version numbers do not track quality on third party benchmarks either: on WildSongBench v5 scores 6.8721, above v6 at 6.5562 and v6 Wild at 6.4195, with v4.5 at 6.6995 and v5.5 at 6.7150, and the vendor itself asks for capability based rather than version based description. Limits: closed with no self-hosting, so unreleased melodies and lyrics must be uploaded; no stable version semantics, meaning a regeneration can sound different after a model update; commercial rights follow the tier; control granularity sits at section and style level with no editable chord track, so theory level edits require Studio multitrack re-arrangement or an open model such as YuE2 that exposes an ABC score. Graded C (vendor claim); the benchmark numbers are submitted by m-a-p and have not been recomputed by us.

6.8721WildSongBench SongBench 均分(v5,第三方测)Vendor Claim · 2026-09
ProductionSunoSite
Suno: the platform that pulls the second half of music production inside one product
Speech & AudioMediaTopC

YuE2-3B: open song generation that exposes the score as an interface

The open song generation model from m-a-p, 3B parameters, weights under CC-BY-NC-4.0, turning lyrics and a style prompt into a complete song with vocals and accompaniment at 48 kHz stereo. We list it as the open-weight top row for music in the audio domain on the strength of a combination that is close to unique among its peers: open weights plus an editable intermediate representation. One AR-NAR Mixture-of-Transformers backbone writes an ABC score (melody and chords) and semantic tokens, flow matching then produces acoustic latents and a VAE decodes them, so the score is an artefact that a person or an agent can read and edit instead of a black box whose only control is another sample. Three cot modes (full, melody, off) map onto composing, covering and direct generation, and the official agentic editing demo runs nine turns across fourteen versions from Mandarin pop to English jazz. Read the benchmark protocol carefully: on 192 WildSongBench prompts the best-of-8 SongBench average of 6.9632 sits above Suno v5 at 6.8721 and Suno v6 at 6.5562, but that is eight candidates with selection against a delivered single candidate; MuLan and AllMusicCaps, the two style-text alignment measures, still favour Suno v5, and PER at 8.44 percent trails Suno v6 Wild at 7.45 and MiniMax Music 3 at 6.27. Cost is the most concrete advantage: a 3.6 minute song in 71 seconds on an RTX 4090 24GB with a peak of 11.18 GiB, and 373 songs per hour on an H800 with vLLM at AR concurrency 32. Limits: non-commercial licence; identity preservation in covers comes almost entirely from a supplied score (CLEWS mAP collapses to 0.006 without one); the benchmark runs on the legacy VAE while the default release is the newer one; no technical report yet and no arena result. Graded C (vendor claim), not recomputed by us.

6.9632WildSongBench SongBench 均分(best-of-8)Vendor Claim · 2026-09
ResearchMultimodal Art Projects (m-a-p)SiteRepo
YuE2-3B: open song generation that exposes the score as an interface