On pretraining
- BERT reads both directions at once
Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.
2 min reads likes comments - Read everything first, specialize later
OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.
2 min reads likes comments