We got back from Google I/O with hands-on access to Omni, Google's new video world model, and we pitted it against Runway Aleph 2 on identical source footage. On the Denoised podcast, we break down what Omni is actually good at (hint: not Veo's job), test the avatar, and work through Flow's agentic editor, Google Pics, and Demis Hassabis on AGI timing.
Quick Take
Google dropped a stack of updates at I/O, and the framing that clicked was simple: Omni is not Veo 4. It is the video version of Nano Banana Pro, a world model built for modifying existing video rather than generating cinematic shots from scratch. That distinction changes how you use it, who it competes with, and why Runway Aleph 2 landed into a much harder comparison.
What We Tested: Omni as a Video-to-Video Model, Not a Veo Replacement
The first instinct on getting Omni access is to compare it to Kling, Seedance, and Veo for cinematic generation. That comparison falls flat. As Joey put it, "this is not Veo 4 and it's a completely different model. This is the video version of Nano Banana Pro."
What Omni is built for:
Video-to-video edits (Addy's preferred term over "editing," which still means cutting in our world)
Inpainting on video with object swaps that preserve motion and shadows
World understanding prompts that reason about objects, physics, and continuity
Sharper text rendering inside generated video
We ran a video-to-video test with a clip of someone walking a dog in Venice. Prompt: turn the dog into a robot. The dog became a matte-plastic robot dog while the leash, the hand grip, the pan, and the background transparency all held. As Addy noted, the model picked a quadruped robot rather than a biped: "that's reasoning. Like it's going through some sort of intelligence."
A second pass turned the dog into a chimpanzee and looked even more grounded once metals and reflections were off the table. A third test on a phone clip of a Google building, prompted "make the building take off like a spaceship," kept the building, added rocket exhaust, and busted out of the ground. We covered Omni's launch in the Gemini app separately for production pros.
The world understanding test: We prompted Omni to create single shots of vintage objects whose first letters spell DENOISED, with each letter visible on the object. Omni produced one clip stringing together a dial phone, newspaper, oil can, ink holder, stopwatch, and an Edison fan. Two letters drifted, but one prompt cut a coherent multi-object sequence with readable letterforms. That is what the "world model" label is pointing at.
What We Tested: The Avatar Feature That Buries Sora's
The avatar tool is buried in the Gemini app UI, and the first round of image-to-video tests with random photos was rough. Likeness fell apart. The actual workflow: calibrate on your phone like a Sora avatar, turn your head left and right, read off numbers. Google ties the avatar to your account.
Addy's take: "I think this is way better than Sora's avatar was." On the champagne shot and the Miami gray t-shirt shot, "that's 98% Joey."
Two caveats:
The model amplifies. Sharper jawline, fuller hair, more jacked. Addy called it the "common denominator handsome" tendency. Google could dial it down if they wanted.
Voice is the weak link. We calibrated in a hotel room on phone audio. The output didn't sound like him. Better mic input likely fixes it, but there is no Adobe Podcast Audio style cleanup happening here.
You can only build one avatar (yourself), and you can't grant permission for others to use it the way Sora characters work. We asked. The answer: maybe later, not on the roadmap. Addy's pitch for the real use case: faceless YouTube channels that want a synthetic talking head. For that audience, the current quality bar is already over the line.
What We Explored: Flow's Agentic Sidebar and a Vibe-Coded Tool Builder
Google Flow picked up the new Omni models, a character system, and a Gemini sidebar that operates as an agentic chat layer over the timeline.
Full breakdown in our coverage of Flow's agent update.
We tested it by asking for "10 different shots of the same person walking in the same field." Flow generated a reference image of the person, generated a reference image of the field, then used both as consistency inputs for the 10 shots. One vague prompt, three steps under the hood.


