Vinson·Li

Essay No. 23

Move 37

AlphaGo beat Lee Sedol four games to one. The move everyone is talking about, how the system found it, and why self-play is the part that matters.


AlphaGo beat Lee Sedol 4-1 yesterday. I predicted in January that it would win, which I mention only because I was expecting something like 3-2 and it wasn’t close until game four.

The moment I’ll remember is move 37 in game two. AlphaGo played a shoulder hit on the fifth line, far from where the fighting was, and the commentators on the English stream thought it was a mistake. Lee Sedol left the room for fifteen minutes. The Korean professionals said nobody would play that. Later in the game it turned out to connect to everything, and AlphaGo won.

DeepMind said afterwards that its policy network had estimated about a 1 in 10,000 chance that a human would play that move. That number comes from the part of AlphaGo that was trained on human games: 30 million positions from strong amateur games on the KGS server, learning to predict what a person would play next. That network is what makes AlphaGo play like a strong human most of the time.

The move came from the other parts. After learning from human games, the policy network played millions of games against earlier versions of itself and was updated with reinforcement learning toward whatever won. A value network was trained on those self-play games to look at a position and estimate who would win. During a real game, a Monte Carlo tree search explores possible continuations, using the policy network to decide which moves are worth looking at and the value network to judge the positions it reaches. The search can end up committing to a move the human-trained intuition considered very unlikely, if the lines that follow from it keep evaluating well.

So move 37 was the system going beyond what it had learned from people. The human data got it started, and self-play plus search took it somewhere humans hadn’t gone. I think that’s the important part of this match, much more than the score. As long as a system only learns to imitate people, it’s capped at roughly human level. Once it can generate its own experience and evaluate it, the cap is gone, at least in a domain with clear rules and a clear winner.

Lee’s win in game four deserves its own mention. His move 78 was a wedge that AlphaGo apparently hadn’t considered seriously, and after it the program played a string of bad moves. That’s the other side of the same coin: when the system’s evaluation is wrong, the search follows it confidently in the wrong direction.

Go is the cleanest possible environment, with perfect information, fixed rules and a winner at the end. The real world has none of those. But the recipe of imitating people to get started, then practicing against yourself with a way of judging the results, seems general enough that I expect to see it again well outside board games.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…