Nobody taught it to walk
DeepMind's locomotion agents learned to run, jump and duck from a reward for moving forward and a varied course. The flailing arms are the point.
DeepMind put out a video last month that I’ve shown to half the people at work. A simulated humanoid runs through an obstacle course, jumping gaps, vaulting walls and ducking under barriers, while its arms windmill in a way nobody would ever design. The paper is “Emergence of Locomotion Behaviours in Rich Environments.”
The setup is what makes it interesting. The agents have simulated bodies (a planar walker, a quadruped, a 3D humanoid) with joints, motors and physics from the MuJoCo simulator. The reward is mostly just forward progress, with small penalties for things like wasted energy. There’s no reward for jumping, no reward for looking like a human, no motion capture data to imitate. The training courses are varied: gaps, hurdles, walls, slopes, platforms of different heights, generated randomly, and the difficulty increases as the agent improves.
What came out is a set of skills nobody specified. The agents learned to jump because jumping was the only way to keep moving forward across a gap. They learned to duck because barriers hit them otherwise. The paper’s argument is that rich environments do a lot of the work that people usually try to put into carefully shaped reward functions. If the world contains many different obstacles, simple goals lead to complex behavior.
The arms are funny, and I think they’re informative too. The reward doesn’t care how the movement looks, only whether the body gets forward. Flailing arms help balance in this simulator and cost almost nothing, so they stay. Humans don’t move like that because we have other constraints: energy, joint limits, pain, evolution tuning us for efficiency over millions of years, and a lifetime of watching other humans. The movement you learn depends on the body you have and what the environment charges you for.
That’s the part I’d push further. These bodies are simplified. Real joints have tight angle limits and torque limits, muscles fatigue, feet slip, and hands (which this humanoid doesn’t really have) have more degrees of freedom than the rest of the body combined. I’d bet that if you gave an agent a body with realistic limits and let it learn from scratch, the movement would get closer to human, because a lot of what makes human movement look human is what our bodies can and can’t do.
And the obvious follow-up, which I’d love to see: start with crawling. A human baby doesn’t go straight to running. It rolls, crawls, pulls itself up, falls a lot, and walks, each step building on the last. A curriculum that follows that order, in a body that follows those limits, might produce much more than locomotion.