Can a network have taste?
Google's NIMA predicts how people would rate a photo, as a distribution, not a single score. Modeling disagreement turns out to be the useful part.
Google Research wrote up NIMA, “Neural Image Assessment,” on their blog last month. It’s a convolutional network that looks at a photo and predicts how people would rate its aesthetic quality. I’ve been meaning to write about it because it touches something I’ve wondered about since the style-transfer paper in 2015: whether taste can be learned at all.
The training data is mostly AVA, a dataset of about 250,000 photos from a photography contest site, where each image was rated from 1 to 10 by a couple of hundred people on average. Earlier work on this problem usually took the mean rating and either regressed it or split photos into “good” and “bad.” NIMA predicts the whole distribution of ratings instead: the probability that a random rater gives it a 1, a 2, and so on up to 10. The loss is the earth mover’s distance between predicted and actual distributions, which respects that a 7 is closer to an 8 than to a 2.
Predicting the distribution sounds like a small technical choice, but I think it’s the key idea. A photo that everyone rates 5 and a photo that half the raters love and half hate can have the same mean. They’re very different photos. The second is probably more interesting, more stylized or more risky. A model that only predicts the mean can’t tell them apart. A model that predicts the spread knows when people disagree, and disagreement is a big part of what taste is.
The results are good enough to be useful. Its rankings correlate reasonably well with human ones, and Google shows it being used to guide automatic photo enhancement, adjusting contrast and brightness toward what the model predicts people will prefer.
I don’t think this is taste in the full sense. It’s learned from one community’s ratings on one site, and it will carry that community’s preferences, maybe for dramatic HDR landscapes and shallow depth of field. It knows nothing about what the photographer meant. But it’s a real, working answer to a question people often assume is unanswerable: can a machine judge whether something looks good? It can, in a statistical, community-specific way, which honestly is also how a lot of human taste works.
The application I care about is creative tools. A lot of creative work is search: try a hundred crops, color grades or edits, keep the best. If a model can score candidates the way a particular audience would, the tool can do a lot of the searching and leave the person to make the final call. I’d expect every photo and video app to have something like this inside within a few years.