49 Atari games from pixels, and zero points in Montezuma's Revenge
DeepMind's DQN paper is in Nature. What the network actually learns, and the game where it learns nothing at all.
DeepMind’s Atari paper is in Nature this week. It’s the full version of the 2013 workshop paper, the one that made Google pay a lot of money for a London lab with no product, and it’s the most interesting thing I’ve read in a while.
The agent sees what a person sees: the raw screen, downscaled to 84 by 84 grayscale, with the last four frames stacked so it can tell which way things are moving. It outputs one of the joystick actions. The only other thing it gets is the score. The same network architecture and the same hyperparameters are used for all 49 games, and on more than half of them it plays at or above the level of a professional human tester.
What the network learns is a Q-function: for the current screen, an estimate of the total future score for each possible action. You play the action with the highest estimate, most of the time, and occasionally a random one so you keep exploring. After each step you nudge the estimate toward “the reward I just got plus the best estimate from the next screen.” That’s textbook Q-learning from the 80s. The reason it didn’t work with neural networks before is that the updates chase themselves and blow up. The paper’s fixes are two engineering tricks. One is experience replay: store past transitions and train on random samples of them, so consecutive frames that look almost identical don’t dominate the updates. The other is a separate target network that’s updated only occasionally, so the thing you’re chasing doesn’t move every step.
The results table is the best part, because it’s honest. At the top are games like Breakout and Video Pinball, where the agent is far better than any person. In Breakout it learned to dig a tunnel through one side of the wall and send the ball behind it, which nobody told it about. At the bottom there’s Montezuma’s Revenge, where it scores zero.
Montezuma’s Revenge is a platformer where you have to climb down ladders, avoid a skull, jump to a rope, and grab a key before you get any points at all. Random button pressing will essentially never do all of that in sequence, so the agent never sees a reward, so it never has anything to learn from. The method needs rewards to show up often enough by accident.
I think this is the more important result, even though it’s a zero. Most of the real world looks more like Montezuma than like Breakout. Nobody gives a baby points for learning to stack blocks, and nobody gives a robot points for the first nine steps of opening a door. A system that only learns when the reward comes to it will stay stuck in those situations. Something has to make the agent go explore on its own before any reward shows up, some reason to try the ladder just to see what’s down there.
Babies do this constantly and it doesn’t look like reward maximization at all. It looks like curiosity. I don’t know how you’d write that down as an objective, but I suspect the next big jump in this area comes from someone who does.