Vinson·Li

Index

On transformer

  1. From embeddings to IDs to what you watch next

    Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you'll watch next.

    3 min
  2. Recommendation as next-token prediction

    A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you'll want next. The item vocabulary finally describes what things are.

    2 min
  3. One network, 604 tasks

    DeepMind's Gato plays Atari, captions images, chats and stacks blocks with a real robot arm, all with the same weights. Turning actions into tokens is the interesting part.

    2 min
  4. Listening sessions are sentences

    Treat each song as a token and each listening session as a sentence, and recommendation starts to look like language modeling. What that framing gets right, and what it misses.

    2 min
  5. One architecture, any input

    DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.

    2 min
  6. An image is worth 16x16 words

    A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.

    2 min
  7. BERT reads both directions at once

    Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.

    2 min
  8. Read everything first, specialize later

    OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.

    2 min
  9. No recurrence, no convolution

    A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.

    2 min