Vinson·Li

Essay No. 47

An agent that dreams its own racetrack

Ha and Schmidhuber's World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I've read this year.


David Ha and Jürgen Schmidhuber put out “World Models” this week, as an interactive web article with demos you can play in the browser. I’ve read it three times and I think it’s the most important thing I’ll read this year.

The agent has three parts.

V, the vision model, is a variational autoencoder. It compresses each 64 by 64 frame of the game into a small latent vector, 32 numbers for the car racing task. The rest of the system never sees pixels again.

M, the memory model, is an RNN with a mixture density output. Given the current latent vector and the action, it predicts a probability distribution over the next latent vector. So it’s a model of how the world changes in response to what you do, but in the compressed space V learned, not in pixels.

C, the controller, is tiny. It’s a single linear layer from the latent vector and M’s hidden state to the action. A few hundred parameters. It’s trained with an evolution strategy (CMA-ES), not backprop, which is feasible because it’s so small.

The heavy lifting is in V and M, which are trained without any reward, just by watching rollouts of random play. The controller only has to learn how to use a representation that already captures what matters and a memory that already knows what tends to happen next.

The part that made me sit up is the dream training. In the VizDoom task, they train the controller entirely inside M’s predictions. The agent never plays the real game during training. It plays the version its own model generates, then gets dropped into the actual game, and it works. There’s a subtle catch the authors are open about: a controller trained in a dream will find and exploit flaws in the dream, like a way to make the fireballs never hit it, that don’t exist in reality. They handle this by raising the temperature of M’s predictions, making the dream more random than reality so exploits don’t pay off.

Why I think this matters beyond a car game and a Doom level: it’s a clean demonstration of learning a model of the world from observation, then using that model to learn behavior cheaply. That’s how I imagine people work. We carry a compressed model of how things behave and do most of our planning inside it. You don’t need to fall off a roof to learn not to walk off it.

The limits are obvious. These worlds are small and simple, and the models are small too. Scaling this to rich 3D physics, with bodies and hands and objects that deform, is a completely different problem. But I expect “learn a world model, then train inside it” to become a major line of research. I also suspect the phrase “world model” is going to show up a lot more in the next few years.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…