Vinson·Li

Index

On video

  1. What would it cost to make Breaking Bad with AI?

    A line-by-line estimate for one 47-minute episode using this year's public video generation prices. Generating the footage is only part of the budget; performance and continuity remain much harder to get right.

    3 min
  2. Taste without a verifier

    Coding agents somehow learned to build good-looking websites, although nothing checks whether a design is beautiful. What video editing can borrow from that: preference models, critics and loops.

    2 min
  3. From embeddings to IDs to what you watch next

    Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you'll watch next.

    3 min
  4. Change everything except the motion

    At I/O this week, YouTube brought Gemini Omni to Shorts Remix: take a Short, change the characters, setting or style, and keep the motion and story. Why remix is the right first experience for video-to-video.

    2 min
  5. Sora is shutting down. Generation isn't a product

    OpenAI closed the Sora app this week, seven months after launching it as a social network, and the API ends in September. What went wrong, and what it says about where generative video belongs.

    2 min
  6. The model plans the shots now

    ByteDance's Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What's left for the editor, and for rights holders.

    2 min
  7. What a video editing agent needs to get right

    I built an agent that edits my social videos end to end using knowledge of my personal taste. What that asks of the timeline, the feedback loop and the product.

    4 min
  8. Sora built a social network

    OpenAI launched Sora 2 as a TikTok-style app where every video is generated. The model is impressive. I'm less sure a feed of generations gives people a reason to keep watching.

    2 min
  9. Genie 3 remembers where you painted the wall

    DeepMind's Genie 3 generates interactive worlds in real time at 720p that stay consistent for minutes. Persistence is the test for a world model, and it just got much better.

    2 min
  10. Sixty episodes, one minute each

    Microdramas are one of the fastest-growing forms of entertainment, and the format exists because of production cost. What happens when AI changes the cost?

    2 min
  11. 62 hours of robot data

    Meta's V-JEPA 2 learns a world model from a million hours of video, then learns to plan robot actions from 62 hours of robot data. Promising, and still missing touch.

    2 min
  12. Veo 3 has sound, and that changes the medium

    Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.

    2 min
  13. What a minute of AI video costs in Beijing vs. San Francisco

    Real per-second prices for Veo 2, Kling 2.0 and Jimeng this month, and why the retake rate matters more than the sticker price.

    2 min
  14. A neural net runs Doom

    Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.

    2 min
  15. Kling came from a short-video company, not a lab

    Kuaishou, TikTok's main rival in China, released a video model that rivals Sora's samples, and ordinary users in China can already try it. Video models get built by whoever has the video.

    2 min
  16. Genie learned to play from videos with no controls

    DeepMind's Genie learned a controllable world model from 2D platformer videos with no action labels, by inferring eight latent actions on its own. The model learns the controls as well as the game.

    2 min
  17. "World simulator" is doing a lot of work

    OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.

    2 min
  18. Four-second clips can't make a movie

    Pika 1.0, Stable Video Diffusion and Runway's latest make beautiful short clips. What's missing is the grammar of film: continuity, screen direction and eyelines.

    2 min
  19. Text to video is next

    Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.

    2 min
  20. It learned Minecraft by watching YouTube

    OpenAI's VPT labeled 70,000 hours of Minecraft videos with the actions players took, using a small model trained on a little labeled data. Passive video became interaction data.

    2 min
  21. Images are solved-ish. Video is where physics lives

    DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.

    2 min
  22. Video is the new floor

    Two weeks into working from home, every interaction is a video call. What that does to expectations for video, and what's missing from the tools.

    2 min
  23. Vine is dead. Short video isn't

    Twitter is shutting down Vine. It didn't lose because six-second videos were a bad idea.

    2 min
  24. A video is not a stack of photos

    Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.

    3 min