Gemini Omni: moving the unit of video delivery from one attempt to a conversational editing process
gemini-omni
Google DeepMind's video and multimodal generation model, officially create anything from any input, starting with video: it edits video through step-by-step conversation, each edit building on the last and keeping the scene coherent, raising the unit of delivery from one shot to footage you can keep revising; the other two surfaces are real-world knowledge and arbitrary reference composition. Read the level with its votes: on the Artificial Analysis arena this site syncs (2026-09-22) gemini-omni-1.1-flash is #1 for text-to-video at Elo 1516 with only 1,784 votes and a ±15 CI, while #2 is the sibling at 1513 with 26,576 votes and ±9, a 3-point gap far smaller than either CI, so statistically indistinguishable: the honest statement is that the Omni family ties for the top. For image-to-video minimax-h3 leads (1495, 57,112 votes) with Omni #2 at 1488 on 3,734 votes. Pricing: $1.50 per million input tokens, $9.00 text and $17.50 video output, 720p about $0.10 per second and a 30-second clip roughly $3. GA to paid-tier developers, none on the free tier. Closed, not self-hostable, not fine-tunable. Graded A (confirmed), our first A-grade video asset; A means the reading is credible and checkable, not undisputed first place.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- Arena T2V Elo(1784 票)
- Confirmed · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it A (confirmed), the first A-grade asset in our video domain, and the basis is specific: capabilities and pricing come from Google's official documentation while the level readings come from a third-party arena (Artificial Analysis) with no stake in this site, and we retain and publish that third party's raw vote counts and confidence intervals. Two independent sources corroborating each other is exactly how our confidence system defines grade A. The meaning of A still needs stating precisely: it says "the readings are credible and the sources are checkable", not "it is the undisputed number one".
The most instructive thing about this asset is how its readings must be read. The T2V leader at 1516 Elo has 1,784 votes and a ±15 interval; second place is the sibling gemini-omni-flash at 1513 with 26,576 votes and ±9. The gap is 3 points and both intervals are far wider than that, so they are statistically indistinguishable. On I2V it is more direct: minimax-h3 leads at 1495 (57,112 votes, ±5) and Omni is second at 1488, 7 points behind with its own ±11 interval - again indistinguishable. The correct statement is "the Omni series sits inside the statistical top band", not "Arena number one". This is the classic shape of a thinly sampled new model: recently listed, few votes, Elo still swinging. We put the vote count in the metric label rather than burying it in the body because an Elo without a vote count is close to misleading in this domain.
On product judgement, Omni's real contribution is not image quality but changing video generation's unit of delivery from one roll of the dice to footage you can keep editing. The official analogy - "Nano Banana, but for video" - is accurate: Nano Banana became the turning point in images because it stayed consistent across a chain of edits, not because a single image looked better. Combined with arbitrary references (image, text, video, audio) resolving into one output, it turns "multimodal input" from two stapled endpoints into a genuinely unified surface. The cost has to be counted too: $17.50 per million video output tokens at 5,792 tokens per second of 720p is about $0.10 per second, so a 30-second clip is roughly $3 and conversational editing costs rounds multiplied by unit price - five iterations lands near $15. A failed roll costs one attempt; conversational editing re-bills the whole output each round. That is a structural difference between two paradigms and the budget must treat them separately. There is also no free tier, so evaluation costs money up front, and being closed with no self-hosting or fine-tuning is a hard gate for teams with data-sovereignty constraints. We have not called the API ourselves, so "physical intuition" and cross-round consistency remain official descriptions rather than readings we measured - the one part of this A-grade asset still hanging open.
The problem it solves: video generation's unit of delivery moves from a single roll of the dice to a conversational editing process
The shared weakness of today's video models is not image quality but controllability: you write a prompt, roll once, and if you dislike the result you rephrase and roll again, with no way to keep the parts you liked and discard only the parts you did not. Google DeepMind positions Gemini Omni as "Create anything from any input - starting with video", using a very precise analogy: "Think of Gemini Omni like Nano Banana, but for video."
Nano Banana became the turning point in image generation not through single-image quality but through staying consistent across a chain of edits. Gemini Omni moves the same property to video: the official description is that you "edit any video through natural, step-by-step conversation", where "every edit you make builds on the one before - maintaining a consistent, coherent scene". That sentence is the whole product. It raises video generation's unit of delivery from "one shot" to "a piece of footage you can keep revising".
Two further official capabilities matter as much:
- Applying real-world knowledge: it combines an intuitive understanding of physics with Gemini's knowledge of history, science and cultural context, which Google frames as "bridging the gap from photorealism to meaningful storytelling". Physical intuition decides whether motion is believable; world knowledge decides whether an instruction like "a Paris street in the 1920s" is executed correctly.
- Referencing anything: any combination of image, text, video and audio references becomes a single cohesive output. That is a genuine multimodal input surface, not two endpoints ("text-to-video" plus "image-to-video") stapled together.
Third-party level: read this first place together with its vote count
On the Artificial Analysis arenas this site syncs (captured 2026-09-22), the Gemini Omni series holds the top two places on text-to-video and second place on image-to-video:
| Board | Rank | Model | Elo | Votes | CI |
|---|---|---|---|---|---|
| Text-to-video (T2V) | 1 | gemini-omni-1.1-flash | 1516 | 1,784 | ±15 |
| 2 | gemini-omni-flash | 1513 | 26,576 | ±9 | |
| 5 | dreamina-seedance-2.0-720p | 1479 | 56,318 | ±8 | |
| 6 | wan3.0 | 1476 | 2,598 | ±13 | |
| Image-to-video (I2V) | 1 | minimax-h3 | 1495 | 57,112 | ±5 |
| 2 | gemini-omni-1.1-flash | 1488 | 3,734 | ±11 | |
| 6 | gemini-omni-flash | 1465 | 87,155 | ±6 |
This site deliberately lists votes and confidence intervals alongside the ranks, because they change how the conclusion should be read:
- The T2V first place has only 1,784 votes with a ±15 interval, while second place is the sibling gemini-omni-flash at 26,576 votes, ±9 and 1513 Elo. The gap is 3 points and both rows' intervals are far wider than 3, so the two are statistically indistinguishable. The correct statement is "the Omni series is tied at the top", not "1.1 Flash is number one".
- The well-sampled high readings belong to other labs: seedance-2.0 has 56,318 votes with a ±8 interval at 1479. Its score sits about 35 points below Omni, and that 35 points is a statistically robust gap.
- Omni does not lead image-to-video: minimax-h3 is first at 1495 (57,112 votes, ±5). Omni 1.1 Flash is second at 1488, 7 points behind, with its own interval at ±11 - a gap smaller than its own uncertainty, again indistinguishable.
Our reading of this asset is therefore: Gemini Omni sits inside the statistical top band of the text-to-video arena without a distinguishable lead, and is likewise indistinguishable from the leader on image-to-video. That differs materially from the colloquial "Arena number one", and the difference is exactly the shape of a new model with few votes - recently listed, thinly sampled, Elo still swinging. We print 1516 as the card reading but put the vote count in the metric label so the premise stays visible on the card itself.
Availability and pricing: promoted from preview to general availability
As of our check, gemini-omni-1.1-flash on the Gemini API (product name Gemini Omni Flash) is generally available to developers on the paid tier, no longer a preview; a separate gemini-omni-flash-preview endpoint still exists alongside it. On the consumer side it is available in the Gemini app and Google Flow.
| Item | Reading |
|---|---|
| Input | $1.50 per 1M tokens (same price for text, image, video and audio) |
| Output (including thinking tokens) | $9.00 per 1M (text), $17.50 per 1M (video) |
| Video billing rate | 5,792 tokens per second of 720p video |
| Effective unit price | About $0.10 per second of 720p video |
| Free tier | Not available |
Put $0.10/second in context: a 30-second 720p clip costs about $3.00, and a conversational editing session with five iterations lands around $15. The real cost of conversational editing is rounds multiplied by unit price. That is the fundamental structural difference from roll-the-dice generation - a failed roll costs you one attempt, while conversational editing re-bills the entire output every round. The absence of a free tier also means evaluation requires payment up front.
Relation to Veo and to the other video assets we catalogue
Google maintains both Veo (positioned as "generate cinematic video with audio") and Gemini Omni. That is a division of labour rather than duplication: Veo produces a high-quality finished clip in one pass, Omni produces multimodal output you can converse with. This site also catalogues Seedance (ByteDance's joint audio-video generation line, delivering 30-second narrative units), Kling, Wan and Flux. Their technical choices differ and the arena readings trade places - the Omni series leads T2V while minimax-h3 leads I2V, and no single model tops both boards. We therefore do not crown one champion in this domain; we catalogue the routes side by side and annotate each with its board reading and vote count.
Boundaries
- The leading readings are thinly sampled: 1,784 votes on T2V and 3,734 on I2V, against tens of thousands to hundreds of thousands for other models on the same boards. Elo swings hard at that sample size, and ranks may well move at our next sync.
- Closed, no self-hosting: no downloadable weights, no fine-tuning, no private deployment. Footage and prompts must pass through Google.
- No free tier: evaluation and iteration both cost money; there is no zero-cost trial path (apart from the consumer Gemini app).
- Multi-round editing cost scales linearly with rounds: as above. This is inherent to the conversational paradigm rather than a defect, but it belongs in the budget.
- "Physical intuition" and "real-world knowledge" are official descriptions, not scorable readings: we have run no physical-consistency benchmark (VideoPhy-style or similar) and measured no cross-round decay in identity or scene consistency.
- Pricing is quoted at the 720p rate: billing rates for higher resolutions are not given in the same place, so cost for long-form or high-resolution work needs separate calculation.
Our verification status
Facts come from three sources read in full: Google DeepMind's official Gemini Omni model page (positioning, the Nano Banana analogy, conversational step-by-step editing, physical intuition and world knowledge, arbitrary reference composition, availability in the Gemini app and Google Flow); the Gemini API pricing page (GA status of gemini-omni-1.1-flash, $1.50 input, $9.00 text / $17.50 video output, 5,792 tokens per second at 720p, roughly $0.10 per second, no free tier, and the coexisting preview endpoint); and this site's own sync of the Artificial Analysis T2V and I2V arenas (captured 2026-09-22, with rank, Elo, votes, confidence interval and preliminary flags checked row by row).
Confidence is graded A (confirmed) - the first A-grade asset in our video domain - on the basis of two independent sources corroborating each other: capabilities and pricing from official documentation, level readings from a third-party arena with no stake in our site, and we retain and publish that third party's raw vote counts and confidence intervals. Grade A here means precisely "the readings are credible and the sources are checkable". It does not mean "it is the undisputed number one" - on the contrary, we state in the body that the top of T2V is statistically indistinguishable from second place, and that it does not lead minimax-h3 on I2V.
What we have not done must be stated just as plainly: we have not called the API to generate video, not measured consistency decay across editing rounds, not run a physical-consistency benchmark, and not done a blind comparison against Seedance, Kling or Veo on identical prompts. Doing so requires paid API quota and a fixed evaluation prompt set - real next work for us. Until then, "conversational editing actually stays consistent" remains an official description rather than a reading we measured.