Every lab is building a playground
DeepMind open-sourced its 3D lab and OpenAI released Universe in the same week. Environments are becoming the new datasets.
Within a few days of each other, the two most visible AI labs released places for agents to learn instead of datasets.
DeepMind open-sourced DeepMind Lab, a 3D environment built on the Quake III engine. An agent sees a first-person view of mazes and rooms, moves around, collects things and solves navigation and memory tasks. It’s where a lot of their recent work on agents that explore 3D spaces was done.
OpenAI released Universe, which is more of a universal adapter. It lets an agent control any program on a computer through a virtual screen, keyboard and mouse, the same way a person would. At launch that’s over a thousand environments: Flash games, browser tasks, and eventually full PC games like GTA V.
A few years ago the fastest way to make progress in AI was to build a bigger labeled dataset. ImageNet did more for computer vision than any single algorithm. The labs seem to have concluded that for agents, the equivalent is environments. If you want a system that learns by acting, you need somewhere for it to act, with enough variety that it can’t memorize the answer, and cheap enough to run millions of times.
I think that’s right, and I think the choice of environment matters as much as the choice of dataset did. An environment determines what can be learned. Atari games teach reflexes and simple strategy. Mazes teach navigation and memory. A web browser teaches following instructions on screens. None of these teach physics in any rich sense: what happens when you push something heavy, how things fall, how soft objects deform, what a hand can do with a cup. Quake physics is not real physics.
The kind of learning I’m most curious about would need something closer to a physics sandbox with a body in it. You’d give an agent limbs with real limits, drop it in a world with objects that behave like objects, and let it play. Robotics simulators like MuJoCo exist, but they’re mostly used for narrow control tasks, not open-ended exploration.
There’s also the question of what the agent wants. Most of these environments come with a score. A child in a playground doesn’t have one. It pokes things to see what happens, and learns a model of the world as a side effect. The environments are finally here. I’m less sure anyone has figured out the right motivation to put in them.