On video
- What would it cost to make Breaking Bad with AI?
A line-by-line estimate for one 47-minute episode using this year's public video generation prices. Generating the footage is only part of the budget; performance and continuity remain much harder to get right.
3 min reads likes comments - Taste without a verifier
Coding agents somehow learned to build good-looking websites, although nothing checks whether a design is beautiful. What video editing can borrow from that: preference models, critics and loops.
2 min reads likes comments - From embeddings to IDs to what you watch next
Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you'll watch next.
3 min reads likes comments - Change everything except the motion
At I/O this week, YouTube brought Gemini Omni to Shorts Remix: take a Short, change the characters, setting or style, and keep the motion and story. Why remix is the right first experience for video-to-video.
2 min reads likes comments - Sora is shutting down. Generation isn't a product
OpenAI closed the Sora app this week, seven months after launching it as a social network, and the API ends in September. What went wrong, and what it says about where generative video belongs.
2 min reads likes comments - The model plans the shots now
ByteDance's Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What's left for the editor, and for rights holders.
2 min reads likes comments - What a video editing agent needs to get right
I built an agent that edits my social videos end to end using knowledge of my personal taste. What that asks of the timeline, the feedback loop and the product.
4 min reads likes comments - Sora built a social network
OpenAI launched Sora 2 as a TikTok-style app where every video is generated. The model is impressive. I'm less sure a feed of generations gives people a reason to keep watching.
2 min reads likes comments - Genie 3 remembers where you painted the wall
DeepMind's Genie 3 generates interactive worlds in real time at 720p that stay consistent for minutes. Persistence is the test for a world model, and it just got much better.
2 min reads likes comments - Sixty episodes, one minute each
Microdramas are one of the fastest-growing forms of entertainment, and the format exists because of production cost. What happens when AI changes the cost?
2 min reads likes comments - 62 hours of robot data
Meta's V-JEPA 2 learns a world model from a million hours of video, then learns to plan robot actions from 62 hours of robot data. Promising, and still missing touch.
2 min reads likes comments - Veo 3 has sound, and that changes the medium
Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.
2 min reads likes comments - What a minute of AI video costs in Beijing vs. San Francisco
Real per-second prices for Veo 2, Kling 2.0 and Jimeng this month, and why the retake rate matters more than the sticker price.
2 min reads likes comments - A neural net runs Doom
Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.
2 min reads likes comments - Kling came from a short-video company, not a lab
Kuaishou, TikTok's main rival in China, released a video model that rivals Sora's samples, and ordinary users in China can already try it. Video models get built by whoever has the video.
2 min reads likes comments - Genie learned to play from videos with no controls
DeepMind's Genie learned a controllable world model from 2D platformer videos with no action labels, by inferring eight latent actions on its own. The model learns the controls as well as the game.
2 min reads likes comments - "World simulator" is doing a lot of work
OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.
2 min reads likes comments - Four-second clips can't make a movie
Pika 1.0, Stable Video Diffusion and Runway's latest make beautiful short clips. What's missing is the grammar of film: continuity, screen direction and eyelines.
2 min reads likes comments - Text to video is next
Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.
2 min reads likes comments - It learned Minecraft by watching YouTube
OpenAI's VPT labeled 70,000 hours of Minecraft videos with the actions players took, using a small model trained on a little labeled data. Passive video became interaction data.
2 min reads likes comments - Images are solved-ish. Video is where physics lives
DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.
2 min reads likes comments - Video is the new floor
Two weeks into working from home, every interaction is a video call. What that does to expectations for video, and what's missing from the tools.
2 min reads likes comments - Vine is dead. Short video isn't
Twitter is shutting down Vine. It didn't lose because six-second videos were a bad idea.
2 min reads likes comments - A video is not a stack of photos
Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.
3 min reads likes comments