Vinson·Li

Essay No. 33

Poker fell

CMU's Libratus beat four top professionals at heads-up no-limit hold'em over twenty days. Hidden information was supposed to protect humans a while longer.


Libratus, a poker program from Carnegie Mellon, finished a twenty-day match in Pittsburgh yesterday against four of the best heads-up no-limit Texas hold’em players in the world. It won by about $1.8 million in chips over 120,000 hands. The pros never had a winning day as a group after the first week.

A year ago, after AlphaGo, the common line was that games like poker would take longer, because Go has perfect information and poker doesn’t. You can’t see your opponent’s cards, bluffing works, and the best play depends on what you think the other player thinks you have. It seemed like the kind of reasoning that needed something more human than a search tree.

It didn’t. From what the CMU team has described, Libratus doesn’t try to read people. It computes an approximation of a game-theoretic equilibrium strategy, one that can’t be exploited on average no matter what the opponent does. The main method is counterfactual regret minimization: the program plays against itself trillions of times, and at every decision point it tracks how much it regrets not having taken each other action, and shifts toward the ones it regrets not taking. Over enough iterations, this converges toward an equilibrium. Bluffing isn’t a special trick in this setup. It falls out of the math, since a strategy that never bluffs is exploitable.

Two details I found clever. First, the full game is far too big to solve, so Libratus solves a simplified version in advance, then re-solves the specific situation it’s actually in, in much finer detail, during the hand. Second, every night of the match it looked at which bet sizes the humans had been using to find gaps in its abstraction, computed better responses for those, and added them to its strategy. The humans were basically finding holes, and the program was patching them overnight.

The broader point for me is that “requires human intuition” keeps turning out to mean “hasn’t been computed yet.” Hidden information, deception and psychology all got handled by self-play and a lot of computation, without any model of the opponent’s mind.

I don’t want to overstate it. Heads-up means one opponent. Six-player poker is much harder, and poker still has clean rules and a clear score. But the space of games where humans are safe keeps shrinking, and the method that keeps winning is the same one: play yourself, measure your regret, repeat.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…