MuZero learns the rules it isn't given
DeepMind's MuZero plays Go, chess, shogi and Atari at top level without being told the rules. It plans inside a model it learned, and the model only predicts what matters.
DeepMind posted MuZero on arXiv last week. It matches AlphaZero at Go, chess and shogi, and sets a new state of the art on the Atari benchmark, and it does this without being given the rules of any of these games.
AlphaZero’s tree search needed a perfect simulator. To look ahead, it asked the real game engine: if I play this move, what’s the next board? That works for board games, where the rules are simple and known. It doesn’t work for Atari, or anything in the real world, where you don’t have a simulator you can query.
MuZero learns its own. It has three functions. A representation function takes the current observation, like the board or the last few Atari frames, and turns it into a hidden state. A dynamics function takes a hidden state and an action and produces the next hidden state and the immediate reward. A prediction function takes a hidden state and outputs a policy and a value. The tree search runs entirely in hidden states: start from the representation of the real observation, then imagine forward with the dynamics function.
The clever part is what the hidden state is trained to capture. It isn’t trained to reconstruct the observation. MuZero never tries to predict the next frame’s pixels. It’s trained so that, after imagining k steps ahead, the predicted rewards, values and policies match what actually happened. So the learned model only needs to represent whatever helps with planning. The color of the background in an Atari game doesn’t matter, so it doesn’t have to be in the state. Nobody forces the hidden state to look like a real board, and it probably doesn’t.
This connects to “World Models” from last year, where Ha and Schmidhuber trained an agent inside its own learned model. That work learned a model that reconstructs observations, then used it. MuZero goes further: it throws out reconstruction and learns a model that’s only good for decisions. I think that’s the direction that scales. Predicting every pixel of the future is enormously expensive and mostly wasted. Most of what’s in an image doesn’t matter for what you should do next.
There’s a question it leaves open that I keep thinking about. MuZero’s model is shaped by the reward. Take away the reward and it doesn’t know what to represent. A child’s model of the world isn’t built for one reward. It’s built before any particular goal exists, and then used for whatever goal comes along. I’d love to see this approach combined with curiosity, a learned model of the world that’s trained to predict whatever’s learnable, and then used for planning toward any goal you give it.