Vinson·Li

Essay No. 20

152 layers, and the trick is learning nothing

Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.


Microsoft Research Asia won this year’s ImageNet competition with a 152-layer network, about 3.6% top-5 error, which is lower than the commonly quoted human estimate. The paper, “Deep Residual Learning for Image Recognition,” went up on arXiv last week. Last year’s winners had around twenty layers. VGG had nineteen.

The problem they solved isn’t the one I would have guessed. You’d think deeper networks fail because they overfit. The paper shows that’s not it: when you take a normal convolutional network and just stack more layers, the training error goes up. A 56-layer network does worse than a 20-layer one on the training set itself. That’s odd, because the deeper network could in principle copy the shallow one and set the extra layers to pass their input through unchanged. The optimizer just can’t find that solution.

Their fix is to make “pass the input through unchanged” the default. Each block of a few layers computes some function F(x), and the block’s output is F(x) + x, via a shortcut connection that skips the layers and adds the input back. If the best thing for a block to do is nothing, it just has to push F toward zero, which is easy. If it has something useful to add, it learns a correction on top of its input.

So each block learns a residual, a small adjustment, instead of a whole new representation. Gradients also flow straight back through the shortcuts, so the early layers of a very deep network still get a useful training signal.

It reminds me of how iterative methods in mechanical simulation work. You don’t solve a nonlinear system in one shot. You start from a guess and apply small corrections until it converges. A residual network looks like a stack of learned correction steps, each nudging the representation a bit.

The practical result is that depth stopped being a problem, and I expect every vision architecture next year to have these shortcuts in it. The more interesting thought is that a lot of the progress in this field comes from making optimization easier, not from new ideas about intelligence. ReLUs, dropout, batch normalization earlier this year, now residual connections. Each one removes a reason training fails.

The same week, a group including Sam Altman and Elon Musk announced OpenAI, a nonprofit research lab with a billion dollars pledged. I don’t know what they’ll work on first. But with results like this coming out every few months, I can see why people want their own lab.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…