Vinson·Li

Index

On architecture

  1. Natively multimodal

    Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.

    2 min
  2. One architecture, any input

    DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.

    2 min
  3. An image is worth 16x16 words

    A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.

    2 min
  4. No recurrence, no convolution

    A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.

    2 min
  5. 152 layers, and the trick is learning nothing

    Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.

    2 min