Stable Diffusion runs on my own computer
Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.
On Monday, Stability AI, together with the CompVis group in Munich and Runway, released the weights of Stable Diffusion. It’s a text-to-image model based on the latent diffusion paper I wrote about in December, trained on a subset of the LAION-5B dataset, and it runs on a consumer GPU with about 10GB of memory. I had it running on my home PC on Tuesday night.
DALL·E 2 and Google’s Imagen are better on some prompts. But they’re behind waitlists and APIs, with content filters and usage terms. Stable Diffusion is a file you download. Five days after release, people have already built plugins for Photoshop and Krita, a dozen web interfaces, image-to-image tools, inpainting, fine-tuning scripts to teach it a specific style or person, and ports that run on Apple Silicon Macs. Nobody at Stability had to plan any of that.
That’s what open weights change. When a capability is behind an API, the company that owns it decides what it’s used for and at what price, and innovation happens at the pace of that company’s roadmap. When the weights are public, thousands of people experiment in parallel, and the useful ideas show up in days. We saw a smaller version of this with open-source machine learning libraries. This is the same thing with the trained models themselves, which are much more expensive to produce.
It also means everything bad that can be done with the model will be done, starting immediately. The license has use restrictions, but there’s no way to enforce them on a model running on someone’s own computer. Non-consensual images of real people, art that imitates living artists’ styles without their consent, and a flood of spam are all happening already. I wrote about deepfakes in 2017 and said the time to design how faces get used was before the capability became easy. It’s easy now, for images in general.
The artist question is the one I find hardest. The model learned from billions of images scraped from the web, including a lot of work by living illustrators who never agreed to it and now see their names used as style prompts. It’s the same consent debt I wrote about with face datasets in 2019, at a much larger scale. I don’t think “it’s legal to train on public data” will be the end of that argument, legally or morally.
My guess is that open models will stay a few months behind the best closed ones on quality, and that this gap won’t matter for most uses. The open ecosystem will win on customization and cost. And whatever happens with images will happen with music and then video. Working in music, I’m thinking hard about what that means for artists.