No recurrence, no convolution
A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.
A paper from Google Brain and Google Research last month has a title that’s a bit of a dare: “Attention Is All You Need.” It introduces an architecture they call the Transformer, which does machine translation with no recurrent layers and no convolutions, and it gets a new best score on English-to-German and English-to-French while training in a fraction of the time. The big model trained in three and a half days on eight GPUs.
Until now, the standard translation model was an RNN encoder and an RNN decoder with attention between them. The RNN reads a sentence one word at a time, carrying a hidden state forward. That’s slow, because each step waits for the previous one, and information from early words has to survive many steps to reach later ones.
Attention on its own works like a soft lookup. Each word produces three vectors: a query, a key and a value. To compute the new representation of a word, you compare its query to the keys of every word in the sentence (a dot product), turn those scores into weights with a softmax, and take a weighted average of the values. So every word looks at every other word directly and pulls in information from the ones that are most relevant to it. In “the animal didn’t cross the street because it was too tired,” the representation of “it” can put most of its weight on “animal” in one step, no matter how far apart they are.
The paper does this several times in parallel with different learned projections, which they call multi-head attention, so different heads can look for different relationships, one for syntax, say, another for which noun a pronoun refers to. It stacks six layers of that plus small feed-forward networks. Since attention has no sense of order, they add position information to each word’s input with sine and cosine patterns at different frequencies.
The speed comes from the fact that all positions are computed at once. That fits GPUs much better than an RNN, which is probably a bigger deal than the translation scores. Faster training means you can train bigger models on more data for the same money.
What I keep thinking about is how little in this architecture is specific to language. It’s a way to let a set of elements exchange information based on content. The elements could be words, but they could also be image patches, notes in a piece of music, frames of video, or the items in someone’s history on a shopping site. My guess is that within a few years we’ll see this block used well outside translation, possibly replacing RNNs for most sequence problems. I’m going to try it on something non-linguistic at work and see what happens.