Vinson·Li

Index

On generative-models

  1. Change everything except the motion

    At I/O this week, YouTube brought Gemini Omni to Shorts Remix: take a Short, change the characters, setting or style, and keep the motion and story. Why remix is the right first experience for video-to-video.

    2 min
  2. Sora is shutting down. Generation isn't a product

    OpenAI closed the Sora app this week, seven months after launching it as a social network, and the API ends in September. What went wrong, and what it says about where generative video belongs.

    2 min
  3. The model plans the shots now

    ByteDance's Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What's left for the editor, and for rights holders.

    2 min
  4. Sora built a social network

    OpenAI launched Sora 2 as a TikTok-style app where every video is generated. The model is impressive. I'm less sure a feed of generations gives people a reason to keep watching.

    2 min
  5. Consistent characters are the unlock

    Google's new image model, known everywhere as Nano Banana, keeps a person looking like themselves across edits and scenes. For AI storytelling, identity preservation matters more than image quality.

    2 min
  6. Sixty episodes, one minute each

    Microdramas are one of the fastest-growing forms of entertainment, and the format exists because of production cost. What happens when AI changes the cost?

    2 min
  7. Veo 3 has sound, and that changes the medium

    Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.

    2 min
  8. What a minute of AI video costs in Beijing vs. San Francisco

    Real per-second prices for Veo 2, Kling 2.0 and Jimeng this month, and why the retake rate matters more than the sticker price.

    2 min
  9. Style is free now. Taste isn't

    GPT-4o's image generation turned the internet into Studio Ghibli for a week. When every style is one prompt away, the scarce thing is judgment.

    2 min
  10. Genie 2 and World Labs in the same week

    DeepMind's Genie 2 turns one image into a playable 3D world, and Fei-Fei Li's World Labs turns one image into a 3D scene you can walk through. Worlds are the next medium after text, images and video.

    2 min
  11. Kling came from a short-video company, not a lab

    Kuaishou, TikTok's main rival in China, released a video model that rivals Sora's samples, and ordinary users in China can already try it. Video models get built by whoever has the video.

    2 min
  12. Music generation's GPT-3 moment

    Suno v3 and Udio generate full songs with vocals and lyrics from a sentence, and some of them are good. What that means for artists, platforms and listeners.

    2 min
  13. "World simulator" is doing a lot of work

    OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.

    2 min
  14. Four-second clips can't make a movie

    Pika 1.0, Stable Video Diffusion and Runway's latest make beautiful short clips. What's missing is the grammar of film: continuity, screen direction and eyelines.

    2 min
  15. The model rewrites your prompt

    DALL·E 3 in ChatGPT doesn't use what you typed. It has a language model write a long, detailed prompt for you. Creative tools are turning into agents that interpret intent.

    2 min
  16. Fake Drake

    An AI-generated song imitating Drake and The Weeknd got millions of plays before it was pulled. Voice is identity, and the industry needs consent and attribution systems, fast.

    2 min
  17. Control beats prompts

    ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.

    2 min
  18. Text to music is a representation problem

    Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.

    2 min
  19. Text to video is next

    Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.

    2 min
  20. Stable Diffusion runs on my own computer

    Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.

    2 min
  21. Images are solved-ish. Video is where physics lives

    DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.

    2 min
  22. Diffusion is going to eat GANs

    Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.

    2 min
  23. "Too dangerous to release"

    OpenAI trained a much bigger language model and is holding back the full version. The unicorn story is impressive. What's actually dangerous is cheap, plausible text at scale.

    2 min
  24. None of these people exist

    Nvidia's StyleGAN generates photographic faces with control over pose, identity and freckles. A face company's view of faces becoming free.

    2 min
  25. An agent that dreams its own racetrack

    Ha and Schmidhuber's World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I've read this year.

    2 min
  26. Face swaps on Reddit: I built the harmless version years ago

    Someone on Reddit is putting celebrities' faces into porn with a home GPU and open-source tools. How it works, and why consent has to be designed in now.

    2 min
  27. 16,000 samples a second

    DeepMind's WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it's so slow.

    2 min
  28. Van Gogh is a Gram matrix

    A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.

    2 min
  29. Two networks arguing

    The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.

    2 min