DeepSeek and the price of intelligence
A Chinese lab released a reasoning model that matches OpenAI's o1 and published how it did it. Nvidia lost about $600 billion in a day. Where the efficiency actually came from.
Nvidia lost about 17% of its market value yesterday, close to $600 billion, the largest one-day loss for any company in history. The trigger was DeepSeek, a Hangzhou lab spun out of a quant hedge fund, whose R1 reasoning model, released a week ago, performs roughly on par with OpenAI’s o1 on math and coding benchmarks, is open-weights under an MIT license, and costs a small fraction as much to use through its API. Its app is at the top of the US App Store.
The market story is “China built frontier AI cheaply, so maybe nobody needs as many GPUs.” I think the technical story is more interesting and points to a different conclusion.
The number people keep quoting is about $5.6 million, from the DeepSeek-V3 paper in December. That’s the cost of the GPU hours for the final training run of V3, the base model, around 2.8 million H800 hours at a rental price. It’s not the total cost of the project. It leaves out research, failed runs, the cost of the GPUs themselves, and the salaries. It’s still a very low number for a model this good, and it comes from real engineering.
V3 is a mixture-of-experts model: 671 billion parameters in total, but only about 37 billion are active for any given token, because a router sends each token to a few specialist sub-networks. You get the knowledge of a huge model at the compute cost of a much smaller one. They also use multi-head latent attention, which compresses the attention cache so long contexts use much less memory, trained in FP8 low precision for much of the network, and wrote a lot of custom communication code to use their export-restricted H800 chips efficiently. None of these ideas is entirely new. Doing all of them together, well, at scale, is.
R1 is the part I find most exciting. Starting from V3, they trained reasoning mostly with reinforcement learning, with rewards that can be checked automatically: whether a math answer is correct, whether code passes tests. Their R1-Zero experiment used almost no supervised examples at all. Through RL alone, the model learned to produce long chains of reasoning, check its own work and backtrack, which they describe as an “aha moment” in training. It’s the same lesson as AlphaGo Zero in 2017 and AlphaProof last summer: when you have a verifier, self-improvement works.
My read on the market reaction: efficiency doesn’t reduce demand for compute, it increases it. When the cost of something useful drops, people use a lot more of it. That’s been true for every computing resource I’ve seen. Cheaper reasoning means reasoning in more products, more agents running longer, and more experiments. And it means the frontier labs will adopt these techniques and spend the same budgets to go further.
The part that should matter to American labs isn’t that DeepSeek is cheap. It’s that a small team under export restrictions published its methods openly and closed the gap in a few months. The moat was never going to be the model weights.