On representation
- Gaussian splatting killed my mesh nostalgia
3D Gaussian Splatting represents a scene as millions of fuzzy, colored blobs and renders it in real time at NeRF quality. It does this without a neural network.
2 min reads likes comments - Predicting in representation space
Meta's I-JEPA is the first concrete result from LeCun's world model agenda. It learns image representations by predicting hidden regions in latent space, with no augmentations and no pixel reconstruction.
2 min reads likes comments - Recommendation as next-token prediction
A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you'll want next. The item vocabulary finally describes what things are.
2 min reads likes comments - Text to music is a representation problem
Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.
2 min reads likes comments - NeRF went from hours to seconds
Nvidia's Instant NGP trains a neural radiance field in seconds with a multiresolution hash table. The representation mattered more than the network.
2 min reads likes comments - Pictures and words in the same space
OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.
2 min reads likes comments - A whole scene stored inside a network
NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.
2 min reads likes comments - BERT reads both directions at once
Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.
2 min reads likes comments