Veo 3 has sound, and that changes the medium
Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.
At I/O on Tuesday, Google announced Veo 3, and the clips people have been posting since are the first generated videos I’ve seen that feel like scenes rather than silent footage. Veo 3 generates audio together with the video: dialogue with lip sync, sound effects that land on the action, ambient sound for the space. An old sailor telling a story on a boat, with the creak of the deck and the sea in the background. A street interview where a character answers a question, in a voice that matches their face. I work at Google, not on Veo, and I’m going by what’s public.
It’s hard to overstate how much sound matters. Film people say audio is half the picture, and anyone who has watched a rough cut without the sound mix knows it’s more than half. Silent generated clips always felt like a mood board. Add dialogue and the characters become people. Add footsteps, doors and room tone and the space becomes real. Add a sound that happens exactly when a glass hits the floor and the physics suddenly feels more believable, even if the visuals are identical.
Generating them together, instead of adding audio afterward with a separate model, matters for the same reason I’ve argued about multimodal models since 2018: modalities explain each other. The sound of an impact is determined by what hit what, how hard, in what room. The lip movement is determined by the phonemes. A model that produces both jointly can keep them consistent in a way a pipeline has trouble with.
Google also launched Flow, a filmmaking tool built around Veo, Imagen and Gemini, with ways to manage characters and scenes and extend shots. That’s the right direction. In 2023 I wrote that the gap between a clip and a film is continuity and scene-level control. Flow is a first attempt at that layer. It’s early, but it’s designed around how filmmakers think.
The other I/O demo I keep thinking about is Gemini Diffusion, an experimental language model that generates text by denoising, like an image model, not left to right one token at a time. It starts with a block of noisy tokens and refines the whole thing in parallel over a few steps. Google says it’s much faster than comparable autoregressive models and competitive on coding benchmarks.
That’s worth pausing on. For five years, “language model” has meant predicting the next token, and “image model” has meant diffusion. If diffusion works well for text, the architectures for every modality start to converge, and generating a scene with dialogue, sound and images could happen in one process, with all of it refined together. It also changes what’s possible for editing: a diffusion model can revise the middle of a passage in context, which is closer to how people actually write.
Taken together, this week suggests the separate tracks of video, audio and text generation are merging. I’d expect the next generation of creative tools to generate complete scenes, sound included, and to be edited more like a film and less like a prompt.