Vinson·Li

Essay No. 113

The model rewrites your prompt

DALL·E 3 in ChatGPT doesn't use what you typed. It has a language model write a long, detailed prompt for you. Creative tools are turning into agents that interpret intent.


DALL·E 3 is now rolling out inside ChatGPT for paying users. You ask for an image in conversation, and ChatGPT generates it. The quality is a big step over DALL·E 2, and the prompt following is the best I’ve used. It gets text in images mostly right, which every earlier model butchered, and it handles “a red cube on a blue sphere, to the left of a green cone” far better.

The detail I find most interesting isn’t the image model. It’s that your words never reach it directly. ChatGPT reads your request and writes a new, much longer prompt, specifying composition, style, lighting and details, and that’s what goes to DALL·E 3. You can see it if you click on the image. I asked for “a cat reading a newspaper in a café” and got a paragraph describing the cat’s breed, the angle of the light through the window, the steam from a coffee cup, and the typography of the newspaper.

OpenAI’s technical note explains why. They trained DALL·E 3 on images with highly detailed synthetic captions generated by a captioning model, instead of the short, noisy alt text that comes with web images. That’s a large part of why it follows prompts better. But it means the model expects long, detailed descriptions, and people write short ones. So a language model sits in between and translates.

That changes what using the tool feels like.

The user’s job moves from writing prompts to expressing intent. “Prompt engineering” as a skill mostly goes away for casual users. You say what you want, possibly vaguely, and the system figures out the specifics, and you react. That’s closer to working with a human illustrator than to operating a tool.

The system becomes an agent with its own taste. When ChatGPT fills in the lighting, the breed and the composition, it’s making creative decisions. Some of them will be good. Some won’t match what you meant. How you steer that, by conversation, by examples, by editing its plan, becomes the core of the interface.

And the pattern generalizes. For video, music, design and code, I’d expect the same architecture: a strong language model that plans and interprets, orchestrating specialized generators and editors that execute. The quality of the result depends on both, but the experience is defined by the planner.

I work on music, so I keep thinking about what this looks like there. “Something for a rainy drive home” is intent. Turning it into a playlist, or eventually into generated music, requires a system that interprets, plans and makes choices. It’s a different kind of product from a search box, and I think it’s where most creative tools are heading.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…