Eleven Music v2.5: the closed-source row that turns music generation into a callable pipeline
eleven-music
The music generation model from ElevenLabs, currently at v2.5, released 2026-09-11 with the release post last updated 2026-09-20. We list it as the engineering-side top row for music in the audio domain, which contrasts with rather than duplicates the Suno entry already catalogued here: the value of Suno concentrates inside the product, a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export, while Eleven Music splits comparable capability into endpoints, namely compose, stream, a structured composition plan, scoring an uploaded video, uploading existing audio, stem separation, finetunes and section level inpainting, plus a marketplace where creators license tracks to each other. Choosing between them therefore needs no audio quality comparison, only one question: is this track something a person sits and adjusts, or something a pipeline requests in volume. The API surface is nine endpoints rather than one generate call, and two details matter for procurement: both compose and stem-separation take a sign_with_c2pa flag that applies to mp3 output so outgoing files can carry content credentials, and output format is tied to subscription tier, with mp3_44100_192 requiring Creator or above and pcm_44100 requiring Pro or above on stem-separation and video-to-music while compose goes up to mp3_48000_320, so the threshold for lossless differs per endpoint. The trap worth memorising is that v2.5 is already the interface default yet the model_id enum of music_v1, music_v2 and music_v2_5 still defaults to music_v1, and output_format=auto resolves per model to mp3_44100_128 on v1 and mp3_48000_192 on v2, so a minimal call that omits model_id silently gets the oldest generation at a lower bitrate; pin both in production code and name mp3_48000_320 when 320kbps is required. The two plan schemas are not interchangeable and the wrong pairing is a hard error: music_v1 takes MusicPrompt while music_v2 and music_v2_5 take CompositionPlan, whose chunks carry a text field with square bracket section names, lyric lines and curly brace inline directions. On the control surface prompt and composition_plan are mutually exclusive, and the request takes music_length_ms from 3000 to 600000, that is three seconds to ten minutes, plus force_instrumental, finetune_id and seed. Rights and commercial use are the most structured part of this line: a multi-year agreement with Universal Music Group was announced alongside v2.5 and the vendor states it is separate from Music 2.5; every track is yours on every plan including Free, Free allows commercial use provided ElevenMusic is credited, lossless downloads are capped at five per day on Free and 400 per month on Pro, tracks built on another artist song through Audio Reference cannot be downloaded, the Marketplace sells licences by usage type with creator earnings starting at 25 percent, and no licence permits distribution to streaming platforms such as Spotify. The official evidence for v2.5 over v2 is a self-run blind test in which v2.5 won the majority of 47,885 paired takes with the widest gap in vocal-led and acoustic-heavy genres. Boundaries: the API is paid-subscription only; vocals are documented for English, Spanish, German and Japanese with no Mandarin, so a Chinese language song is more practical on YuE2 or Suno; seed does not guarantee reproducibility; and it is closed source with no weights, so the self-hostable alternatives here are YuE2 for music and VoiceStudio or Kokoro-82M for speech. Graded C (vendor claim): quality was not recomputed and we ran no blind listening test, but the API contract is verifiable documentation and every clause of it is listed in the body.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- API 单曲时长上限(music_length_ms 3000-600000)
- Vendor Claim · 2026-09
- MATURITY
- Production
- research → demo → product → production
Our takeWe grade it C (vendor-claim) because the only quantitative quality evidence we can cite is a vendor-run blind test: the same prompt rendered as two takes and compared as pairs, with v2.5 preferred in the majority of 47,885 pairs and the widest gap in vocal-led and acoustic-heavy genres. The sample size is respectable, but it is a first-party A/B with no third party benchmark reading to cross-check, and we ran no blind listening of our own, so the contract puts it at C. What does hold up at grade B is a different thing: the API contract is verifiable documentation.
music_length_msfrom 3000 to 600000, threemodel_idvalues withmusic_v1as the default,output_format=autoresolving tomp3_44100_128on v1 models andmp3_48000_192on v2 models, the video-to-music ceilings of 10 files and 200MB and 600 seconds, and the tier gates on stem-separation where 192kbps mp3 needs Creator whilepcm_44100needs Pro. None of that is evidence about sound. It is the boundary a reader can write code against, and listing every one of those in the body is precisely how a C row stays useful.Its role at the top of the audio ladder is the engineering-side representative for music, which contrasts with rather than duplicates Suno, our other closed-source music row. Suno pulls the back half of music production into one product, a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export, and its value is a person finishing a song in an interface. Eleven Music splits comparable capability into endpoints, compose, detailed, stream, composition-plan, video-to-music, upload, stem-separation, finetunes and inpainting, and its value is a program finishing a song. Choosing between them therefore does not require comparing audio quality, only asking one question: is this track something a person sits and adjusts, or something a pipeline requests in volume. Automatic video scoring, podcast intros, batch game music and an agent scoring its own output are the second case; an independent musician reworking a chorus is the first.
The third reason has nothing to do with music and matters most to procurement: this is the line where rights are documented in the most structured way. A multi-year strategic agreement with Universal Music Group spanning licensing and product development, which the vendor states is separate from Music 2.5. Ownership that attaches to the track, with terms changes applying only to new tracks. Commercial use on Free provided ElevenMusic is credited. Downloads blocked on tracks built from another artist song. A marketplace that sells licences in four usage types, Social Media, Paid Marketing, Offline and Enterprise, with creator earnings starting at 25 percent. A prohibition, under every licence, on distributing to streaming platforms such as Spotify and Apple Music. And a
sign_with_c2paflag on both compose and stem-separation. When an enterprise adopts generative music the thing that actually goes wrong is the rights chain and content provenance, not a 3dB quality difference, and this line turns the rights chain into terms you can read. That is scarce among closed-source vendors.Four practical boundaries. First, do not trust SDK defaults: v2.5 is the default in the product interface while the API default for
model_idis stillmusic_v1, and combined withoutput_format=autothat yields the oldest model at a lower bitrate with nothing in the response to flag it. Pin both the model id and the output format in production code. Second, the plan schemas are not interchangeable:MusicPromptfeeds v1 only andCompositionPlanfeeds v2 and v2_5 only, so the wrong pairing is a hard error, and section durations can be relaxed on v1 withrespect_sections_durations=falsewhile v2 models always enforce them. Third, Mandarin vocals are not in the documented set, which lists English, Spanish, German and Japanese; for a Chinese-language song move to YuE2 or Suno rather than forcing it through a prompt. Fourth, the docs lag the release: the product page still names v2 as the interface default and still describes 44.1kHz exports at 128 to 192kbps, while the release post from the same week says v2.5 is default and that every plan including Free gets lossless downloads. Where they disagree, the parameter enums in the API reference are the authority.
The problem it solves: turning song generation into a music pipeline a program can call
Eleven Music is the music generation model from ElevenLabs, currently at v2.5, released 2026-09-11 (release post last updated 2026-09-20). Its division of labour with Suno, the other closed-source top row for music in the audio domain, is not about sound quality but about programmability. The value of Suno concentrates inside the product: a Studio multi-track timeline, Custom Models, up to twelve stems and MIDI export. Eleven Music splits the same set of capabilities into endpoints: compose, stream, a structured composition plan, scoring an uploaded video, uploading existing audio, stem separation, finetunes, section level inpainting, plus a marketplace where creators license tracks to each other. That is why we list it as the engineering-side top row for music in the audio domain: if the job is to embed music generation into your own product or agent workflow, this is the most complete closed-source surface we have found.
The official claim for v2.5 over v2 is better audio quality and better prompt adherence with no change in capability set, so Audio Reference, composition plans and inpainting all carry over. The evidence offered is a self-run blind test: the same prompt rendered as two takes, one from each model, compared as pairs, with v2.5 preferred in the majority of 47,885 pairs. The gap is widest in vocal-led and acoustic-heavy genres, named as R&B, soul, hip hop, rock, metal, orchestral and cinematic. That is a vendor A/B rather than a third party leaderboard and we did not recompute it, which is what sets the confidence grade for this row.
The API surface: nine endpoints, not one generate call
| Endpoint | Method and path | What it is for |
|---|---|---|
| Compose music | POST /v1/music | One song from a prompt or a composition plan, returned as an audio file stream |
| Compose music with details | POST /v1/music/detailed | Same generation with a structured response carrying the song id and metadata that inpainting needs |
| Stream music | POST /v1/music/stream (plus a detailed variant) | Audio while it generates, for interactive products and for Music as a Flows node |
| Create composition plan | POST /v1/music/composition-plan | Expands one prompt into a JSON plan you can edit before spending a generation on it |
| Video to Music | POST /v1/music/video-to-music | Score uploaded video: up to 10 files concatenated in order, 200MB combined, 600 seconds total, with an optional description under 1000 characters and up to 10 style tags |
| Upload Music / Stem Separation | POST /v1/music/upload, POST /v1/music/stem-separation | Split existing audio into stems and return a ZIP; the docs warn explicitly about high latency on long files |
| Finetunes | /v1/music/finetunes (list, create, get and friends) | Train a private style variant on your own original audio and pass finetune_id at generation time |
Two details are easy to miss and matter for procurement. First, C2PA signing: both compose and stem-separation take a sign_with_c2pa flag that applies to mp3 output, so outgoing files can carry content credentials. Where AI-generated music has to survive a platform review step, that is a hard requirement rather than a nice extra. Second, output format is tied to subscription tier: on stem-separation and video-to-music the docs state that mp3_44100_192 requires Creator or above and pcm_44100 requires Pro or above, while compose goes up to mp3_48000_320. Getting lossless therefore has a different threshold on each endpoint, so check per endpoint instead of assuming one product-wide rule.
The control surface: the composition plan is the most valuable part of this product line
Prompt-only generation exposes music_length_ms (3000 to 600000, that is 3 seconds to 10 minutes; omit it and the model picks a length from the prompt), force_instrumental (true guarantees an instrumental), finetune_id, store_for_inpainting and seed. prompt and composition_plan are mutually exclusive, and seed and force_instrumental only work with a prompt.
Delivery-grade arrangement goes through a plan, and there is one thing to read carefully: the two plan schemas are not interchangeable and using the wrong one is a hard error.
music_v1 takes MusicPrompt | music_v2 and music_v2_5 take CompositionPlan | |
|---|---|---|
| Structural unit | sections (SongSection) | chunks |
| Lyrics | A separate lines field, at most 30 lines per section and 200 characters per line | Written into text: a section name in square brackets such as [Verse 1], then lyric lines, with inline directions in curly braces such as {scratching}; same 30 line and 200 character limits |
| Style | Global positive_global_styles and negative_global_styles plus per-section positive and negative local styles | Each chunk carries positive_styles and negative_styles, and the docs note that the styles of the first chunk matter most because they set the tone and genre of the whole song |
| Section timing | duration_ms 3000 to 120000, and respect_sections_durations=false relaxes it so the model can adjust individual sections for quality and latency while preserving the total | duration_ms 3000 to 120000, always enforced, and that flag is ignored |
| Inpainting | Section level source_from with song_id, range and negative_ranges | Same semantics through the inpainting guide |
Two further controls live in the product rather than in the request body. Audio Reference takes an uploaded clip of roughly 30 seconds or less to guide a v2 or v2.5 generation, and every upload is screened for copyright compliance. The official wording is precise about what it does: it influences sound, production style, instrumentation, tempo and mood, and it does not copy or remix the uploaded audio. Finetunes come in two kinds, curated ones pretrained by ElevenLabs across global genres and custom ones trained on your own original audio. Together those two answer the question a prompt cannot, which is how to make your music keep sounding like yours.
The trap worth memorising: the API default model is still music_v1
In the product interface v2.5 is already the default for prompted and reference generation, but the model_id enum on POST /v1/music is music_v1, music_v2 and music_v2_5 with a default of music_v1. At the same time output_format=auto means pick the best format for the selected model, which resolves to mp3_44100_128 for v1 models and mp3_48000_192 for v2 models. Put together, a minimal call that omits model_id gets the oldest generation of the model at a lower bitrate, and nothing in the response says so. The official quickstart does pass model_id="music_v2_5" explicitly, yet plenty of real integrations get carried along by an SDK default. Make two things explicit in production code: the model id, and the output format, naming mp3_48000_320 when you need 320kbps instead of trusting auto.
A second caution is that the documentation lags the release. The product docs page still states that Music v2 is the default model in the Eleven Music interface while the 2026-09-11 release post states that v2.5 is now the default. On export, the product docs describe MP3 at 44.1kHz and 128 to 192kbps with other formats coming soon, while the same week post announces lossless downloads on every plan including Free. Where those conflict our rule is to trust the parameter enums in the API reference, because that is the surface the SDKs and the validation layer actually consume.
Rights and commercial use: a UMG agreement, tiered licences, and a marketplace starting at 25 percent
For enterprise procurement the primary risk in music generation is rights rather than audio quality, and this is the closed-source line where rights are documented in the most structured way.
- A multi-year strategic agreement with Universal Music Group was announced alongside v2.5 and spans licensing and product development. The vendor states explicitly that this deal is separate from Music 2.5, so do not read a new model as a licensed catalogue.
- Ownership and commercial use: every track you make is yours on every plan including Free. Free allows commercial use provided ElevenMusic is credited. Lossless downloads are capped at five per day on Free and 400 per month on Pro. The permissions attach to the track, so cancelling or downgrading does not retroactively change what you may do with tracks already created, and any future tightening of terms applies only to tracks made from that day onwards.
- One hard block: tracks built on another artist song, meaning an Audio Reference to a third party work, cannot be downloaded. The vendor frames this as symmetric protection, the same rules that protect your music protect theirs.
- Music Marketplace: creators publish songs they generated with ElevenLabs and buyers purchase a licence by usage type, which are Social Media (your own channels including monetized video, podcasts, personal sites), Paid Marketing (paid ads, client and agency work, commercial products), Offline (live events, trade shows, physical venues and installations) and Enterprise and Custom (TV, film, cinema, VOD and OTT streaming, large-scale distribution, radio, large studio games). Creator earnings start at 25 percent of the purchase price and pay out through the existing ElevenLabs payout system.
- Two things are prohibited under every licence: distributing tracks to music streaming platforms such as Spotify, Apple Music or SoundCloud, and reselling, sublicensing or claiming ownership of the music.
Boundaries and failure modes
Six of them. First, the API is paid-subscription only, so the free tier can create in the web product but cannot drive a backend batch job. Second, the vocal language set is narrow: the documented languages are English, Spanish, German and Japanese, with no Mandarin vocals, so for a Chinese-language song the more practical rows in our audio domain are YuE2 (open weights, non-commercial licence) or Suno. Third, seed does not guarantee reproducibility: the official wording is that the same seed with the same parameters helps consistency but exact reproducibility is not guaranteed and outputs may change across system updates, and seed cannot be combined with prompt anyway. Projects that need a stable long-term sound should store the plan and the finished audio rather than replay parameters. Fourth, ten minutes is the ceiling for a single track, and going longer means segmenting and reassembling with inpainting or stems, which is a different workflow and a different latency budget. Fifth, the quality evidence is a vendor-run blind test with no verifiable third party benchmark reading, which is why confidence stays at C. Sixth, it cannot be self-hosted: closed source with no weights, so unreleased melodies and lyrics travel through their service, and data-sovereignty-sensitive work has to be assessed on that basis. The self-hostable alternatives in our audio domain are YuE2 for music and VoiceStudio or Kokoro-82M for speech.