On diffusion
- A neural net runs Doom
Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.
2 min reads likes comments - Control beats prompts
ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.
2 min reads likes comments - Text to video is next
Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.
2 min reads likes comments - Stable Diffusion runs on my own computer
Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.
2 min reads likes comments - Images are solved-ish. Video is where physics lives
DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.
2 min reads likes comments - Diffusion is going to eat GANs
Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.
2 min reads likes comments