LeCun's path, read carefully
Yann LeCun's position paper argues that intelligence needs world models that predict in representation space, not pixels. Where I agree, and where I'd push back: bodies and hands.
Yann LeCun posted a long position paper on OpenReview last week, “A Path Towards Autonomous Machine Intelligence.” It’s about sixty pages, not a results paper, and it lays out how he thinks machines could learn and reason more like animals and people. I spent the weekend with it. A lot of it overlaps with things I’ve been writing about here for years, so I want to go through it carefully.
The starting observation is one I agree with completely. A teenager learns to drive in about twenty hours. A baby learns that unsupported objects fall within a few months, mostly by watching and poking. Our best AI systems need enormous amounts of data or trial and error to learn much less. LeCun’s explanation is that people and animals learn world models, internal models of how the world works, mostly through observation, and then use them to plan and to learn new tasks quickly.
The architecture he proposes has several modules: perception, a world model, a cost module (built-in drives plus a trainable critic), an actor that proposes actions, short-term memory, and a configurator that sets up the others for the current task. The world model is at the center.
The key technical idea is the Joint Embedding Predictive Architecture, JEPA. Here’s the distinction he draws. A generative model predicts the future observation itself, like the next video frame, pixel by pixel. But most of the detail in the next frame is unpredictable and irrelevant: exactly how leaves move, the texture of the carpet. Trying to predict it wastes capacity and produces blurry averages. A JEPA encodes both the current observation and the future observation into representations, and predicts the representation of the future from the representation of the present. The encoders are free to leave out details that can’t be predicted. Prediction happens in an abstract space, where only what matters needs to be modeled.
The obvious danger is collapse: if both encoders output a constant, prediction is trivially perfect. So the paper spends a lot of time on training methods that keep the representations informative, like VICReg, which his group published last year.
This connects to MuZero, which I wrote about in 2019, learning a model that only predicts what matters for planning, and to the curiosity work that uses learned feature spaces to ignore noise. LeCun’s version is more ambitious: a hierarchy of these predictors, operating at different time scales, learned mainly from observation without rewards.
Where I’d push back is on how much can be learned from watching. The paper leans heavily on video. Babies do learn a lot by watching, but the most important things they learn about physics, like weight, grip, balance and what their own body can do, come from acting and feeling. Touch, proprioception and force aren’t in any video. I think a world model that’s supposed to be grounded in the physical world needs a body with those senses, with realistic limits, and a lot of time spent using it, the way a child spends a year learning to stand. Hands especially. A lot of what we understand about objects, we learned by holding them.
That’s a friendly disagreement about emphasis. The core idea, predicting in representation space rather than generating every detail, seems right to me, and I suspect it’ll be one of the more important directions of the next few years.