Vinson·Li

Essay No. 95

We've been undertraining

DeepMind's Chinchilla paper says large language models have far too many parameters for the data they see. The fix is more data, and that raises a question about where it comes from.


DeepMind posted “Training Compute-Optimal Large Language Models” on Tuesday, and I think it’s going to change how everyone trains these models.

The question is simple. If you have a fixed compute budget, how should you split it between making the model bigger and training it on more data? The influential scaling laws from OpenAI in 2020 suggested that, as compute grows, you should mostly grow the model and grow the data more slowly. That’s roughly what the field did. GPT-3 has 175 billion parameters and was trained on about 300 billion tokens. DeepMind’s own Gopher has 280 billion parameters, also on about 300 billion tokens.

DeepMind trained over 400 models of different sizes on different amounts of data and fit the curves again. Their conclusion is that model size and training tokens should grow in roughly equal proportion. For a given budget, the optimal model is much smaller and sees much more data than people have been doing. The rough rule of thumb from the results is about 20 tokens per parameter.

To test it, they trained Chinchilla, 70 billion parameters on 1.4 trillion tokens, using the same compute as Gopher. Chinchilla beats Gopher, GPT-3 and other larger models on almost every benchmark. It’s also four times smaller, so it’s much cheaper to run.

The practical consequence is obvious: the big models of the last two years were undertrained, and the next generation will be smaller for their compute and fed a lot more text. The less obvious consequence is the one I keep thinking about. If a 70-billion-parameter model wants 1.4 trillion tokens, a model ten times bigger, trained optimally, wants something like 14 trillion. There may not be that much high-quality text in the world that can be used. Estimates vary a lot, but the books, papers, Wikipedia and good web pages add up to a finite number, and we might hit it within a few years.

So where do the next trillion tokens come from? I see a few possibilities. Lower-quality web text, which has diminishing returns. Other languages, which are underused. Code. Transcribed audio and video, which is a lot of language. And, the one I find most interesting, data that models generate themselves, through interaction, like AlphaZero’s self-play, if we can figure out how to judge its quality.

Beyond text, the richest source of data that nobody has really tapped is the physical world: video of everything that happens, and eventually experience from agents acting in it. Text is a small, heavily compressed record of what people noticed. The world itself is much larger. My guess is that the models that eventually go furthest won’t be trained mostly on text.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…