A neural net runs Doom
Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.
A paper from Google Research and Tel Aviv University this week is called “Diffusion Models Are Real-Time Game Engines,” and it means it literally. GameNGen runs Doom, interactively, at about 20 frames per second on a single TPU, with no game engine. Every frame is generated by a neural network from the previous frames and the player’s inputs. You move, shoot, open doors, pick up items, and the model draws what should happen next. In a short test, human raters could only tell real gameplay clips from generated ones slightly better than chance. I wasn’t involved in this work; this post is based on the public paper.
The training has two stages.
First, a reinforcement learning agent plays Doom a lot, and every frame and action is recorded. The agent isn’t trying to be good. It’s there to generate varied gameplay, the way a playtester would.
Second, a Stable Diffusion 1.4 model is fine-tuned to predict the next frame, conditioned on a window of previous frames and the sequence of actions. Actions are embedded and replace the text conditioning.
The hard problem is drift. When you generate autoregressively, feeding each generated frame back as input, small errors accumulate and within seconds the image degrades into mush. Their fix is to add noise to the context frames during training, at random levels, and tell the model how much noise was added. So the model learns to correct corrupted inputs instead of trusting them, and at inference it can recover from its own small mistakes.
The results have limits. Memory is short, only about three seconds of past frames, so the game state has to be inferred from what’s on screen. Health and ammo show on the HUD, so they persist. Things off screen that you haven’t seen for a while can change. And Doom is a single fixed game the model saw an enormous amount of.
But the idea is important. This is what I’ve been describing as a world model since the 2018 “World Models” paper: a learned function from state and action to next state, good enough that you can interact with it. It’s Genie from March, but at a much higher fidelity, on a real 3D game. And it points to something I find exciting: games and simulations that are learned, not programmed. You wouldn’t write a physics engine, you’d train one from footage.
The implication I keep coming back to runs the other way. If a model can learn a game engine from recorded play, can it learn the real world from recorded video with actions attached? Robot data, driving data, first-person video of people doing things. A learned simulator of the real world, good enough to practice in, is what embodied agents need to learn at scale. This is a toy version of that, and it works.