One architecture, any input
DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.
DeepMind posted “Perceiver: General Perception with Iterative Attention” last week. It takes on the most annoying practical problem with Transformers: attention compares every input element with every other one, so the cost grows with the square of the input length. That’s fine for a sentence of a few hundred words. It’s a disaster for raw inputs like an image with 50,000 pixels, a second of audio with 48,000 samples, or a video.
Vision Transformers dealt with it by cutting images into patches, which works but is a design choice specific to images. Every modality ends up with its own tokenizer and its own tricks.
The Perceiver’s fix is to add a bottleneck. It keeps a small array of learned latent vectors, say 512 of them. Instead of the input attending to itself, the latents attend to the input through cross-attention: each latent can pull information from any of the input elements, which costs latents times inputs, not inputs squared. Then the latents run through ordinary Transformer layers among themselves, which is cheap because there are only 512. That repeats a few times, with the latents going back to the raw input to pick up more detail.
The input can be anything that’s a set of elements with position information: pixels, audio samples, points from a lidar scan, or all of them at once. The paper uses essentially the same architecture on ImageNet from raw pixels, on AudioSet audio and video, and on 3D point clouds, and gets competitive results on each, without assuming anything about the structure of images or audio beyond the position encodings.
I find the design pleasing because it matches how I’d want a perception system in a body to work. The world sends you a huge, messy stream from many senses at once. You don’t process every pixel against every other pixel. You keep a compact internal state and keep going back to the senses to query what you need. The latent array is a kind of working memory, and cross-attention is looking at the world with a question in mind.
The practical implication is more modest but important. If one architecture handles every modality, you can train one model on all of them together, and let attention find relationships between, say, the sound of an impact and the frame where it happens. Separate modality-specific models that get glued together at the end can’t do that as naturally.
I’d expect this idea, a small latent state that reads from large multimodal inputs, to show up in a lot of places, including robotics, where the inputs are cameras, joint sensors, touch and sound all at once.