Vinson·Li

Essay No. 59

The Bitter Lesson, read by someone who hand-built facial features

Rich Sutton says 70 years of AI research show that general methods plus computation beat human knowledge every time. He's right, and it stings. My one objection is about data.


Rich Sutton published a short essay last week called “The Bitter Lesson.” It’s about a page and a half. The argument is that the biggest lesson from 70 years of AI research is that general methods that use more computation win in the long run, and methods that build in human knowledge lose. Chess got solved by search, not by encoding what grandmasters know. Go got solved by search and learning. Speech recognition moved from hand-built models of phonemes to statistical methods to deep learning. Computer vision moved from hand-designed features to convolutional networks trained on data. Each time, researchers tried to build in their understanding of the problem, it helped for a while, and then something general with more compute beat it.

He calls the lesson bitter because researchers keep having to relearn it, and because building in your own knowledge is satisfying and feels like progress.

I read it and recognized myself immediately. My master’s thesis was a hand-built model of faces. I chose landmarks based on anatomy. I designed a parameterization so that the jaw and nose were controlled separately. I spent months deciding which radial basis functions to use so the lips wouldn’t stick together. All of that is exactly what Sutton says loses. And it did lose: by 2016, networks trained on lots of photos were fitting faces better than anything I’d designed, and last year StyleGAN learned a coarse-to-fine face hierarchy on its own from 70,000 photos.

So I agree with him. Mostly.

My objection, or maybe my addition, is about what the computation is spent on. Search and learning need something to search over and something to learn from. For Go, the simulator is perfect and free, so self-play generates unlimited experience. For vision, the web handed us billions of images. The general methods won in places where data or experience was abundant. Where it isn’t abundant, like robotics or anything that happens in the physical world, general methods are stuck waiting.

That’s why I think the next version of this argument is about interaction. The general method that will win the physical world is probably an agent that generates its own experience by acting: poking, grasping, falling, trying again, in simulation and in reality. The human knowledge we’ll be tempted to build in is things like hand-designed grasps, scripted motions and physics rules. The lesson predicts those lose to agents that learn them from experience. The bottleneck is building environments, bodies and motivation so the agent can collect that experience at scale.

One place I think structure survives is the body itself. Sutton’s argument is about knowledge in the algorithm. The shape of the body and the senses aren’t knowledge in that sense. They determine what experience is even possible. A hand with touch sensors can learn things a gripper with a camera never will, however much compute you throw at it.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…