One network, 604 tasks
DeepMind's Gato plays Atari, captions images, chats and stacks blocks with a real robot arm, all with the same weights. Turning actions into tokens is the interesting part.
DeepMind published “A Generalist Agent” on Thursday. The agent, Gato, is one Transformer with 1.2 billion parameters, and with the same set of weights it plays Atari games, captions images, chats, and controls a real robot arm to stack blocks, across 604 tasks in all.
The trick that makes this possible is serialization. Every kind of data is turned into a flat sequence of tokens from one shared vocabulary. Text is tokenized the normal way. Images are cut into 16 by 16 patches, like the Vision Transformer. Continuous values, like joint angles and torques for the robot, are scaled and discretized into 1,024 bins, each bin a token. Discrete actions, like Atari button presses, are tokens too. An episode of a task becomes one long sequence: observation tokens, then action tokens, then the next observation, and so on. The model is trained to predict the next token, just like GPT, except many of the tokens are actions.
At test time, you give it a prompt that shows which task it’s doing, usually a demonstration of the task, and it produces action tokens, which get decoded back into button presses or joint commands.
A lot of reactions this week have been either “this is AGI” or “this is just multitask learning.” I think both miss what’s interesting.
It’s not general intelligence in any meaningful sense. Gato is mostly good at each task because it was trained on expert data for that task. It doesn’t show much transfer, where learning one task makes it much better at an unrelated one. And at 1.2 billion parameters it’s small, partly because it has to run the robot arm in real time. On many tasks it’s worse than a specialist.
What I think matters is the demonstration that actions can be treated as just another kind of token in the same sequence as images and text. That’s a clean interface. It means the whole machinery built for language models, including scaling, pretraining on huge data and prompting, can in principle be applied to acting in the world. If you can put a robot’s camera frames, joint positions, touch readings and the instruction “put the cup in the sink” in one sequence, a big enough model trained on enough of those sequences might learn to act.
The obstacle is data. Language models have the internet. Nobody has an internet of robot actions. Most of Gato’s control data came from reinforcement learning agents trained in simulation. To get somewhere with real bodies, we’ll need either huge amounts of robot experience, which is slow and expensive to collect, or a way to learn actions from the enormous amount of video of people doing things, where actions aren’t labeled at all. That second option seems like the more promising problem to work on.