Vinson·Li

Essay No. 98

It learned Minecraft by watching YouTube

OpenAI's VPT labeled 70,000 hours of Minecraft videos with the actions players took, using a small model trained on a little labeled data. Passive video became interaction data.


Last month, writing about Gato, I said the most promising problem in robotics might be learning actions from the huge amount of video of people doing things, where actions aren’t labeled. On Thursday OpenAI published a clean demonstration of that idea: Video PreTraining, or VPT, in Minecraft.

The problem with learning to act from YouTube is that a video shows what happened, but not what the player pressed. You see the character walk forward and swing at a tree, but the keyboard and mouse inputs aren’t recorded. Imitation learning needs those actions.

VPT’s solution has three steps.

First, they paid contractors to play Minecraft for about 2,000 hours while recording both the screen and every keyboard and mouse action. That’s a small labeled dataset.

Second, they trained an inverse dynamics model on it. Given a stretch of video frames, including frames after the moment in question, predict what action was taken at that moment. That’s a much easier problem than predicting what a player will do next, because you can see the outcome. If the view turned left, the mouse moved left. The model only needs to learn the mapping between effects and actions.

Third, they ran the inverse dynamics model over about 70,000 hours of Minecraft gameplay from the internet and labeled every frame with a guessed action. Then they trained a policy on all of it by behavioral cloning: given the frames so far, predict the next action, like a language model predicting the next word.

The resulting agent can chop trees, make planks and a crafting table, and swim, hunt and pillar-jump. After fine-tuning with reinforcement learning, it crafted a diamond pickaxe, which takes a skilled person about twenty minutes and around 24,000 actions. That’s the first time an agent has done it using the normal human interface.

The general idea is what matters. A small amount of labeled interaction teaches you to read actions off video. Then all the unlabeled video in the world becomes interaction data. Minecraft is an easy case because the actions are keyboard and mouse. For real-world video, of people cooking, fixing things, walking, playing sports, the “actions” are the movements of a human body, and inferring them from video means estimating 3D pose, hand configurations and forces. Harder, but the same shape of problem.

If that works, the internet’s video becomes to robotics what its text became to language models. There’s far more video of people doing physical things than there will ever be robot data. The missing piece is a body that maps onto a human one well enough to use what it learns from watching, which is one more argument for humanoids with real hands.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…