LeCun leaves, Marble ships
In one week World Labs launched Marble, Yann LeCun confirmed he's leaving Meta to start a world-model company, and Google shipped Gemini 3. The world-model bet went public from three directions.
It’s been quite a week for the idea I’ve been writing about on this blog for years.
On November 12, World Labs launched Marble, its first commercial product: a generative world model that creates persistent 3D worlds from text, images, video or a rough 3D layout, which you can explore, edit and export. It’s the productized version of the demo I wrote about last December, and it’s much bigger in scope. You can expand a world, combine worlds, and export them as Gaussian splats or meshes for use in other tools.
On Wednesday Yann LeCun confirmed he’s leaving Meta, where he has run FAIR and been chief AI scientist for over a decade, to start a company focused on what he calls Advanced Machine Intelligence: systems that understand the physical world, have persistent memory, can reason and plan. In other words, the agenda from his 2022 position paper, JEPA and world models, as a company.
The day before, Google released Gemini 3, which I won’t say much about since I work at Google, except that it’s a big step on multimodal understanding, which is a piece of the same picture.
Three years ago, “world model” was a phrase you mostly heard in reinforcement learning papers. Now it’s what major labs, a well-funded startup led by Fei-Fei Li, and one of the founders of deep learning are all betting their next years on. It’s worth being clear that they mean different things by it.
World Labs means a model that generates explicit, persistent 3D space. Its output is a world you can look at and walk through, stored in a form graphics tools understand. It’s grounded in geometry. The emphasis is on spatial intelligence: what’s where, what shape it is, how it looks from anywhere.
LeCun means a model that predicts in abstract representation space how the world will change, especially in response to actions, to support planning. It doesn’t generate the world at all. It predicts the relevant parts of what comes next. V-JEPA 2 in June was the clearest example so far.
The frame-generating models like Genie 3 mean a third thing: a simulator that renders what you’d see next given what you do, with the world existing only in the predictions.
In my Genie 2 post last December I guessed the answer would be a hybrid: explicit structure for what must persist, learned dynamics for what changes. I still think that, and I’d add one thing. None of the three, as they exist today, has a body. Marble’s worlds are places a camera moves through. V-JEPA 2’s robot has a gripper and no touch. Genie’s agents are cameras with a few actions. The world models that matter most for physical intelligence will need an agent inside them with realistic physical limits: mass, joints, balance, hands. That’s the piece I’d most like to see someone take on next.