Vinson·Li

Essay No. 11

The dress is a world-model bug

Half the internet sees white and gold, half sees blue and black. The disagreement is about lighting, and it says something about how vision works.


Everyone I know spent Thursday night arguing about a dress. Half of them see white and gold, half see blue and black, and each side thinks the other is joking. I see blue and black, for the record, and our studio’s designer saw white and gold with total confidence until she opened it in Photoshop.

The pixels themselves aren’t really in dispute. If you sample them, the “blue” parts are a light, washed-out periwinkle and the “black” parts are a brownish gold. So nobody is seeing the pixel colors directly. Everyone is seeing an interpretation of them.

The reason is color constancy. A white shirt reflects bluish light in the shade and yellowish light under indoor bulbs, and you still see it as white in both places, because your visual system estimates the lighting and discounts it. That estimate is a guess, built from the context in the image and a lifetime of experience with how light usually behaves. It is almost always right, which is why we don’t notice it happening.

The dress photo is overexposed and cropped tight, so it gives very few cues about the lighting. It could be a blue and black dress under warm, yellowish light, or a white and gold dress in bluish shadow. Both explanations fit the pixels. Your brain picks one and then shows you the colors that would follow from it, with no indication that it made a choice. Some vision scientists are already suggesting that people who spend more time in daylight or artificial light may lean different ways, but I haven’t seen data yet.

What I take from it: perception is inference. You don’t measure the world, you run a model of how it probably is and check that the model is consistent with what comes in through your eyes. When the input is ambiguous, the model decides, and you experience the decision as a plain fact.

This matters for computer vision more than it might seem. A network trained on labeled photos learns to go from pixels to labels. It doesn’t have an explicit notion of “the lighting in this scene” that it’s reasoning about. It just learns whatever correlations get the label right. That works until you give it the dress, or a car in a lighting condition that wasn’t in the training set. People fail on the dress too, but they fail in an interesting, structured way, because there’s a model underneath with a specific assumption you can point at.

I spend my days trying to make 3D faces look alive, and this is a helpful reminder. The viewer isn’t grading my pixels. Their brain is building its own model of a face, lighting and all, and my job is to not give it anything that contradicts the model.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…