Vinson·Li

Index

On deep-learning

  1. An image is worth 16x16 words

    A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.

    2 min
  2. A whole scene stored inside a network

    NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.

    2 min
  3. The embedding table is the model

    Facebook open-sourced DLRM, its deep learning recommendation model. It shows what big recommenders actually look like: mostly memory, with a small network on top.

    2 min
  4. Read everything first, specialize later

    OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.

    2 min
  5. Faces help you hear

    Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.

    2 min
  6. Can a network have taste?

    Google's NIMA predicts how people would rate a photo, as a distribution, not a single score. Modeling disagreement turns out to be the useful part.

    2 min
  7. No recurrence, no convolution

    A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.

    2 min
  8. Google Translate got better overnight, and Chinese speakers noticed first

    Neural machine translation replaced phrase tables, starting with Chinese to English. Notes from someone who reads both, and why the zero-shot result is the bigger story.

    2 min
  9. 152 layers, and the trick is learning nothing

    Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.

    2 min
  10. Van Gogh is a Gram matrix

    A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.

    2 min
  11. DeepDream sees dogs everywhere

    Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.

    2 min
  12. 49 Atari games from pixels, and zero points in Montezuma's Revenge

    DeepMind's DQN paper is in Nature. What the network actually learns, and the game where it learns nothing at all.

    2 min
  13. Two networks arguing

    The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.

    2 min
  14. A computer captioned a photo. It didn't see the photo

    Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.

    2 min
  15. A video is not a stack of photos

    Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.

    3 min
  16. Graduated. My thesis in one paragraph, and how deep learning will eat it

    What my master's thesis on 3D face meshes does, which part of it I think neural networks will replace soon, and which part I think they won't.

    2 min