Vinson·Li

Essay No. 25

Google built its own chip for neural nets

The Tensor Processing Unit has been running in Google's data centers for a year. When a company designs custom silicon for a workload, the workload is here to stay.


At I/O this week Sundar Pichai mentioned, almost in passing, that Google has built its own chip for machine learning. It’s called the Tensor Processing Unit. They say it has been running in their data centers for more than a year, that it powers RankBrain in search and Street View processing, that AlphaGo used it in the Lee Sedol match, and that it gives “an order of magnitude better-optimized performance per watt” for machine learning than the alternatives.

There aren’t many technical details yet. It’s a custom ASIC that fits in a hard drive slot in their servers, and it’s built for running trained networks, not training them. The general idea is easy to guess. A neural network at inference time is mostly large matrix multiplications, and you don’t need 32-bit floating point precision for them. A chip that only does low-precision multiply-and-add, in huge parallel arrays, can skip most of what a CPU or even a GPU spends transistors on.

What I find more interesting than the chip is what it says about Google’s expectations. Designing a custom chip costs a lot of money and takes years, and once it’s fabricated you can’t change it. Companies only do that for workloads they’re sure will still be large when the chip arrives. The TPU project must have started around 2013, when deep learning had just won ImageNet and speech recognition and was still new in production. Google committed silicon to it then. That’s a strong signal about how much of their computing they expect to be neural networks.

I think the same logic will spread. Nvidia already sells GPUs to deep learning researchers as a major market, not a side effect of gaming. Phones will get neural network accelerators, because running a vision model on the CPU burns the battery, and every company that does face or photo features on device wants that problem gone. We fight exactly that problem in client apps today: a model that works fine on a desktop GPU turns a phone into a hand warmer.

For a studio like ours, the practical point is that the cost of running models is going to drop fast, on servers and on phones. Features that are too expensive to run today, like analyzing every video frame on device, will become cheap enough to be defaults within a few years. When we scope products with clients now, I’d rather design for where the hardware will be when the app ships than for where it is today.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…