Vinson·Li

Essay No. 96

Images are solved-ish. Video is where physics lives

DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.


OpenAI announced DALL·E 2 last Wednesday, and the images are a big step from anything I’ve seen. “An astronaut riding a horse in photorealistic style.” “A bowl of soup that is a portal to another dimension, as digital art.” Realistic lighting, coherent compositions, plausible textures, in many styles. It can also edit existing images from a text instruction, adding a flamingo to a pool with the right reflections.

In December I predicted 2022 would be the year of image generation. It’s April and I think that’s already true.

The method, which OpenAI calls unCLIP, chains together pieces I’ve written about. A text prompt is encoded with CLIP’s text encoder. A “prior” model, itself a diffusion model, maps that text embedding to a CLIP image embedding: a representation of what an image matching the text should contain. Then a diffusion decoder generates an image from that image embedding, with upsamplers bringing it to 1024 by 1024. Generating through CLIP’s image space means the model works with semantics first and pixels second, which is why it’s good at following prompts and varying style.

It also fails in ways that are informative. It can’t write text in images; you get letter-shaped gibberish. It struggles with binding attributes to objects, like “a red cube on top of a blue cube,” and often swaps them. It can’t count reliably. These are all things where CLIP’s representation is weak, and the decoder can’t recover what the embedding doesn’t contain.

The question I’ve been asking since Wednesday: what about video? It’s the obvious next step, and everyone will try it. I think it’s much harder than going from GANs to DALL·E 2, for a specific reason.

An image generator needs to know what things look like. A video generator also needs to know how things behave. When a glass falls, it has to accelerate at the right rate, hit the floor, and shatter into pieces that stay pieces. When a person walks, their feet have to contact the ground, their weight has to shift, their clothes have to follow. When the camera moves, objects must keep a consistent 3D shape and position. Every frame has to be consistent with physics and with the frames before it. You can hide a lot of physical nonsense in a single beautiful image. You can’t hide it across two hundred frames.

So I think video generation will force these models to learn something much closer to a model of the physical world, or to fail visibly until they do. The first text-to-video models, which I’d guess we’ll see this year, will be short, low resolution and physically wrong in amusing ways. How fast they get physically right will say a lot about whether generative models can learn physics from watching.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…