Vinson·Li

Essay No. 42

AlphaGo Zero threw away human games and got better

Starting from random play, with no human data, it beat the version that beat Lee Sedol 100 games to 0 after three days. Human knowledge was a ceiling.


DeepMind published AlphaGo Zero in Nature this week. It starts with only the rules of Go. No human games, no hand-crafted features, nothing but random play at the beginning. After three days of playing itself, it beat the version that beat Lee Sedol, 100 games to 0. After 40 days it beat the Master version that swept Ke Jie in May.

When I wrote about Move 37 last year, I described AlphaGo as two stages: learn from human games to get started, then improve through self-play. I said the human data got it started and self-play took it past people. The first stage wasn’t needed, and was probably holding it back.

The architecture also got simpler. The original had separate policy and value networks and used fast random rollouts to evaluate positions. Zero has a single residual network with two heads, one giving move probabilities and one estimating who will win. There are no rollouts. The training loop is clean: at every move of a self-play game, a tree search guided by the current network produces a better move distribution than the raw network alone. The network is then trained to predict that improved distribution and the eventual winner. So search makes the network better, and the better network makes the search better. It ran on four TPUs, versus the forty-eight used in the Lee Sedol match.

What I find most striking is what happened to human Go knowledge along the way. DeepMind shows Zero discovering well-known joseki, standard corner sequences, in the order humans discovered them over centuries, then abandoning some of them for variations people hadn’t played. It reinvented a lot of what people knew and then kept going.

I’ve been saying for a while on this blog that the most interesting systems learn from their own experience instead of from our labels. This is the strongest evidence yet. In a domain with a perfect simulator and a clear win condition, human data isn’t a head start. It’s a bias, and the system does better without it.

The catch is how special Go is. You can simulate a game of Go perfectly and cheaply, and you always know who won. Most things we care about don’t have that. You can’t simulate the real world perfectly, and “who won” is rarely obvious. So I think the lesson isn’t “throw away human data everywhere.” It’s that the most valuable thing you can build for a problem is a good simulator and a good evaluator. If you have both, self-play will probably beat whatever humans can teach it.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…