BERT reads both directions at once
Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.
Google AI released BERT this month, “Bidirectional Encoder Representations from Transformers,” and it broke most of the language understanding leaderboards in one go. On SQuAD, a reading comprehension benchmark, it scores above the human baseline. In June I wrote that pretraining was coming for language the way ImageNet came for vision. BERT is that, at a larger scale and with one important change in the objective.
OpenAI’s GPT, which I wrote about in June, was trained to predict the next word. That means each word’s representation can only depend on the words to its left, because the model can’t be allowed to see what it’s predicting. For generating text, that’s fine. For understanding a sentence, it’s a strange handicap. The meaning of “bank” in “I sat on the bank of the river” depends on words that come after it.
BERT’s fix is the masked language model. During pretraining, 15% of the words in each input are hidden, and the model has to predict them from everything else in the sentence, on both sides. Since the hidden words are hidden, the model can look in both directions without cheating. They add a second task, predicting whether sentence B actually follows sentence A in the original text, to teach it something about relationships between sentences. The large model is a 24-layer Transformer encoder with about 340 million parameters, pretrained on Wikipedia and a books corpus.
It’s easy to lump BERT in with text generation because it’s “a language model,” but I think that’s a mistake. BERT doesn’t generate text in any natural way. What it produces is a representation: for every word in the input, a vector that encodes what that word means in this particular context. Then for a task you put a small layer on top and fine-tune. Classify a sentence, find the answer span in a paragraph, tag names. The value is in the vectors.
That distinction between a model that represents and a model that generates will matter more over time, I think. Representation models are for understanding and retrieval: search, classification, matching a query to a document, a user to an item. Generative models are for producing things. They share architecture and a lot of training data, but they’re trained with different objectives and good at different things.
For us at Amanda, the representation idea is familiar. A face recognition model is a representation model too. It maps a face to a vector so that the same person ends up close together, and it never generates anything. BERT is doing something similar for words in context. I expect the same approach to reach almost anything that has structure: code, music, product catalogs, user histories. I’d especially like to see someone treat a user’s sequence of clicks the way BERT treats a sentence.