Curiosity is a prediction error
A Berkeley paper gives agents an internal reward for being surprised by their own predictions. It learns Mario with no score at all, and it answers a question I asked two years ago.
Two years ago, writing about DeepMind’s Atari agent scoring zero in Montezuma’s Revenge, I said the missing piece was curiosity and that I didn’t know how you’d write it down as an objective. A paper from Berkeley last month does exactly that: “Curiosity-driven Exploration by Self-supervised Prediction,” by Deepak Pathak and colleagues.
The idea is that the agent should be rewarded for encountering things it can’t predict. It keeps a forward model that takes the current state and the action it’s about to take and predicts the next state. After acting, it compares the prediction to what actually happened. The size of the error becomes a reward, called an intrinsic reward because it comes from inside the agent, not from the game score. So the agent is pushed toward situations where its model of the world is still wrong, and as the model improves, those situations stop being rewarding and it moves on to new ones.
The obvious problem is noise. If the forward model tries to predict raw pixels, then leaves blowing in the wind or static on a TV are endlessly unpredictable, and a curious agent would stare at them forever. The paper’s fix is the clever part. It doesn’t predict pixels. It learns a feature space using an inverse model: given two consecutive states, predict which action the agent took. The features that help with that task are the ones related to things the agent can influence or that influence the agent. Blowing leaves don’t help you guess the action, so they don’t show up in the features. The forward model predicts in that feature space, and curiosity is measured there.
The results are what got my attention. In VizDoom it learns to navigate mazes with very sparse rewards. In Super Mario Bros, with no game reward at all, curiosity alone takes it through a large part of the first level. It learns to kill enemies and jump over pits because dying sends it back to the start, and the start is boring now.
That’s much closer to how I imagine children learn. Nobody scores a baby for dropping a spoon off a high chair. It does it because it isn’t sure yet what will happen, and it keeps doing it until it is. What gets learned along the way, gravity, that spoons make noise, that parents pick them up, is a model of the world that becomes useful later for things the baby hasn’t thought of yet.
I think this is one of the most important ideas in reinforcement learning right now, even if the benchmarks are small. Rewards from the environment are rare in the real world. A system that generates its own reward from wanting to understand things can learn in far more places.