On architecture
- Natively multimodal
Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.
2 min reads likes comments - One architecture, any input
DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.
2 min reads likes comments - An image is worth 16x16 words
A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.
2 min reads likes comments - No recurrence, no convolution
A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.
2 min reads likes comments - 152 layers, and the trick is learning nothing
Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.
2 min reads likes comments