Gemini Omni generates 4K video with synchronised audio and voiceover in the same pass, and can hold one character with one voice across many clips. The tests below were filmed in May 2026 — since then it has reached the Gemini API, and Veo 4 has arrived as a direct competitor. Both change the recommendation.
Most AI video models generate a silent clip and leave you to add sound. Gemini Omni writes the picture and the audio together in one pass, including spoken voiceover, which changes what a single generation is worth.
This article covers what the filmed tests actually showed, what it costs, where it fails, and — because three months is a long time in this field — which of the video's conclusions no longer hold.
Want to try the character workflow? Rangy has Gemini Omni built in, with a saved library of your characters and voices.
Try it freeWhat is Gemini Omni?
Google's video generation model, launched at I/O in May 2026. It produces clips at up to 4K, accepts images, video and audio as input, edits existing footage, and can build a multi-shot sequence from a single prompt rather than one continuous take.
The multi-shot behaviour is the part that surprises people. Give it a product and a rough brief and it will write its own ad scenario, cut between angles, and score the result — work that would normally be three separate tools and an edit.
In the filmed tests it handled cinematic product ads for ice cream, shampoo, sunglasses, a watch and skincare well enough that several were described as ready to run. An educational time-lapse and vertical social clips also held up. The weak case was 3D, covered further down.
Does it really generate the audio too?
Yes, and this is the genuine differentiator. Sound effects, ambience and spoken voiceover are all generated in the same pass as the picture and arrive synchronised to it, rather than as a separate text-to-speech track you have to align afterwards in an editor.
The practical consequence is that a finished thirty-second ad can come out of one generation instead of a generation plus a voice tool plus an edit. For social content, where the audio bar is low and the turnaround is short, that collapses most of the pipeline.
It is also the reason the per-clip price looks high next to silent models. You are not comparing like with like — a silent clip still needs sound before it can be used.
How do you keep the same character and voice?
By registering the character once instead of describing it every time. You supply a reference image and a description, attach a voice profile, and the model returns a character you can then select for any subsequent clip — so the face and the voice both carry across scenes.
The workflow is four steps: generate a character image, create a voice profile, combine the two into a saved character, then pick that character from a dropdown in any later generation. Once it exists, the same person with the same voice appears in every clip without being re-described.
This is the part of the video that has aged best, because it solves a problem prompting cannot. A character described in words drifts; a registered character does not. It is the same principle as the reference-image approach in the character consistency guide, applied to voice as well as face.
Where the library lives matters. A character is only useful if it persists. In Rangy the voices and characters you create are stored locally as a library file and offered as a dropdown on any Gemini Omni generation, so a brand mascot built once in June is still selectable in December.
Where can you actually use it?
Directly through Google's own API now, which was not true when the tests were filmed. At the time Omni was a consumer product with no developer access, so the video routes around it through a third-party provider. That workaround is no longer the only option.
This is the one conclusion in the video worth correcting rather than repeating. It launched at I/O in May as a consumer feature; developer access followed at the end of June. If you are reading the original test and wondering why it never mentions going direct, that is why.
The routes available now:
- Google directly, through the Gemini API and AI Studio. The obvious choice if you are already building on Google's stack.
- A third-party provider such as Kie, which is what the video demonstrates and what desktop apps route through.
- Inside a desktop app that wraps the API and adds the character library, so you are not managing raw calls.
None of these is obviously right. It is a workflow decision now rather than an availability constraint, which is a much better position to be choosing from.
What does Gemini Omni cost?
Roughly seventy-five cents for an eight-second clip through a third-party provider, which works out near nine cents a second. Google publishes its own per-second rate for direct API access, and it has changed since launch — check their current pricing rather than any figure quoted in a video, including this one.
That is expensive next to silent video models, where eight seconds runs from about twelve cents to a dollar depending on the model and resolution. The comparison is misleading, though, because those clips arrive without sound.
| Model | ~8 seconds | Audio included |
|---|---|---|
| Gemini Omni | ~$0.75 | Yes, with voiceover |
| Kling 3 | ~$0.56, or ~$0.80 with sound | Optional |
| MiniMax H3 | ~$0.90 at 768P | Native stereo |
| Seedance 2 Mini | ~$0.50 at 720p with a reference | No |
Rates from Rangy's live pricing tables, checked 15 August 2026. Provider rates change, and Google's direct API pricing is set separately.
Judged as a finished shot with sound rather than as raw footage, Omni is priced about where you would expect. Judged per second of silent video, it looks costly.
Where does it fall down?
Three-dimensional scenes with objects that have to be tracked. In the filmed test a handheld object visibly switched hands mid-clip — the model kept the scene coherent and lost the object's continuity. Body movement in character work is also softer than dedicated animation models manage.
Both limits point the same way: Omni is strong at atmosphere, framing and sound, and weaker at rigid physical continuity. That is a reasonable trade for advertising and social content, and a poor one for anything where an object's position carries the meaning.
- Object tracking in 3D is the clearest failure. Test any clip where something is held, passed or moved before committing to it.
- Full-body animation is not its strength. The character workflow is built for a recognisable presenter, not for choreography.
- Anything that must match reality exactly — a real product, a real place — needs a reference image, and the same caution as in the garment fidelity test applies.
Should you use Omni or Veo?
Veo 4 did not exist when these tests were filmed and now overlaps directly — native 4K, longer clips, character consistency across cuts and multi-speaker audio. The honest answer today is that they are close enough to test both on your own brief rather than to take a recommendation.
What can be said without hand-waving is that Omni's distinguishing feature is the registered character with a locked voice, reusable across arbitrary clips. If your work is a recurring presenter, mascot or spokesperson, that is the capability to weigh.
If your work is one-off cinematic shots, the newer competition has closed much of the gap, and the sensible move is a side-by-side on a brief you actually have. Broader comparisons of the video field are in the guide to AI video generators compared.
When should you use something else?
When you do not need the audio, when the clip must show a real product or place accurately, when physical continuity matters, or when you are generating at volume. A silent model at a third of the price is the better tool for anything you were going to score yourself.
Stated plainly, because this guide is published by a company that sells access to the model:
- You already have sound. Licensed music, a real voiceover, or an existing edit. Then you are paying for a feature you will discard.
- Volume testing. Twenty ad variants at seventy-five cents each is fifteen dollars; the same twenty at silent-model rates is closer to ten, and you only need sound on the winner.
- Accuracy to a real subject. Generated video is not documentation, whatever the resolution.
- Choreography or physical action. Object tracking is the documented weak point.
- You are on Google's stack already. Then go direct rather than through any intermediary, including ours.
Frequently asked questions
Does Gemini Omni generate audio and voiceover?
Yes. Sound effects, ambience and spoken voiceover are produced in the same pass as the picture and synchronised to it, rather than added as a separate track afterwards. That is its main differentiator against silent video models, and the reason a per-clip price comparison against them is not like for like.
Can you use Gemini Omni through an API?
Yes, though not at launch. It arrived as a consumer feature at Google I/O in May 2026 and reached the Gemini API and AI Studio at the end of June, which is after the tests in this article were filmed. Older guides that tell you a third-party provider is the only route are describing the situation before that change.
How much does Gemini Omni cost per clip?
About seventy-five cents for eight seconds through a third-party provider, roughly nine cents a second. Google prices direct API access separately and has revised it since launch, so check their current pricing page rather than a figure quoted in an article or video. Remember the price includes generated audio.
How do you keep the same character across clips?
Register the character rather than describing it. Supply a reference image and a description, attach a voice profile, and the model returns a saved character you select for later generations. Face and voice then persist across scenes without being re-prompted, which is what descriptions alone cannot achieve.
What is Gemini Omni bad at?
Physical continuity in three-dimensional scenes. In the filmed test an object being handled visibly switched hands mid-clip. Full-body animation is also softer than dedicated animation models produce. It is strong on atmosphere, framing and sound, and weaker wherever an object's exact position has to hold across frames.
Is Gemini Omni better than Veo?
They overlap heavily now. Veo 4 shipped after these tests were filmed with native 4K, longer clips, character consistency across cuts and multi-speaker audio. Omni's distinguishing feature remains the registered character with a locked voice reusable across arbitrary clips. For anything else, test both on a real brief rather than trusting a ranking.
Can Gemini Omni edit existing video?
It accepts images, video and audio as inputs and can edit footage rather than only generating from text, which is unusual among video models. In practice that makes it useful for extending or restyling material you already have, though the same caution about physical continuity applies to edited clips as to generated ones.
The bottom line
Use Gemini Omni when the clip needs sound and a recurring character, and something cheaper when it does not. The registered character with a locked voice is the capability worth paying for; the 4K and the ad-writing are increasingly available elsewhere.
The wider lesson from re-checking a three-month-old test is worth more than the model verdict. Two of the video's framing assumptions — that you could not get Omni directly, and that nothing else did this — were true when it was filmed and are not true now.
That is not a criticism of the test. It is how fast this moves, and why the useful habit is to check availability and pricing yourself before acting on any comparison, this one included.
"This software has increased my workflow speed tenfold, and the output quality it has delivered in my work has been exceptional."
Gemini Omni with a character library
Rangy runs Gemini Omni alongside Seedance, Kling and MiniMax on your own API key — and keeps the voices and characters you build in a local library, selectable on any later generation.
Download Rangy Free Or watch the full test →Mac & Windows · Free plan, no credit card · 5 generations a day
The capability findings come from two filmed tests: a model walkthrough published 21 May 2026 covering a dozen-plus examples including product ads, a time-lapse, vertical social clips and the 3D failure case, and a character-and-voice workflow published 15 June 2026. Both are embedded above. Availability was re-checked against Google's own channels on 15 August 2026 rather than taken from the videos, which is how the correction in the access section arose — Omni was a consumer-only feature when the first test was filmed and reached the Gemini API and AI Studio at the end of June. No direct Google per-second price is quoted here, because it is set by Google, has been revised since launch, and a stale figure would be worse than none; the seventy-five cents for eight seconds is the third-party rate shown in Rangy's own interface and was re-checked on 15 August 2026. Competing model rates come from the same live pricing tables.
This guide is published by Rangy, which sells access to this model and its competitors, so it is not a neutral source. It also tells you to go direct to Google if you are already on that stack, which is in when to use something else.