Vinson·Li

Essay No. 51

A robot hand with a hundred years of practice

OpenAI's Dactyl learned to rotate a block with a human-like robot hand, entirely in simulation. Why hands, and why randomizing the simulator is the clever part.


OpenAI published Dactyl yesterday. It’s a Shadow Dexterous Hand, a robot hand built to resemble a human one with 24 degrees of freedom, holding a block with letters on its faces. You give it a target orientation and it rotates the block in its palm, using its fingers to roll and pivot it, until the right face is up. It does this dozens of times in a row.

The policy was trained entirely in simulation, then run on the real hand without further training. The simulated training amounts to about a hundred years of experience, done in about fifty hours on thousands of CPU cores. It uses the same reinforcement learning algorithm OpenAI used for its Dota bots.

The hard part of going from simulation to reality is that simulators are wrong. Friction, the stiffness of the fingers, the exact size and weight of the block, the delay between command and motion: none of it matches the real robot exactly, and a policy trained in one precise simulation usually fails on real hardware. Dactyl’s answer is domain randomization. In every simulated episode, they randomize the physics: friction, masses, object size, actuator strength, the noise in observations, even the lighting and colors seen by the cameras. The policy never sees the same world twice, so it can’t learn to depend on any particular setting. It has to learn a strategy that works across all of them, and the real world is just one more variation.

What I love about the result is that the hand discovered grasps people use. Finger gaiting, where you walk the object around with alternating fingers, and pivoting, and using gravity. Nobody programmed those. They came out of the physics and the hand’s shape.

I’ve thought for a long time that hands are underrated in AI. Most robot manipulation uses parallel-jaw grippers, which are basically pliers. They’re reliable and simple and they make grasping a solvable engineering problem, but they rule out almost everything interesting a hand does. A human hand has more degrees of freedom than the rest of the arm, and much of what we call dexterity and tool use, maybe much of what we call intelligence, comes from having one. If you want an agent to learn how the physical world works by handling it, you need to give it something to handle it with.

The limitation that stands out is touch. The Shadow hand has touch sensors, but Dactyl didn’t use them. It relied on cameras to track the block and joint sensors for the fingers. People manipulate objects mostly by feel. I suspect the next big step will be policies that learn from touch as well as sight, and simulators good enough to model contact at the level fingertips experience it. That’s a hard simulation problem, and a very physical one.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…