DeepDream sees dogs everywhere
Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.
Google Research posted “Inceptionism” last week, and the images from it are everywhere now. Clouds that turn into fish, skies full of pagodas, and above all dogs. Dog faces in trees, in plates of spaghetti, in the background of the Mona Lisa.
The technique is simple once you see it. Normally you feed an image into a trained classifier and adjust the network’s weights to make the output match the label. Here the weights are frozen and you adjust the image instead. Pick a layer, and change the input pixels by gradient ascent so that whatever that layer is detecting gets stronger. Whatever the network faintly sees in the image, it now sees more of, and the change feeds back into the next iteration. Run it repeatedly, zooming in a bit each time, and you get the fractal dream images.
Lower layers respond to edges and textures, so optimizing them gives you strokes and swirls, almost like a painting filter. Higher layers respond to object parts and whole objects, so optimizing them grows eyes, snouts, wheels and buildings wherever something vaguely resembles one.
The dogs are the most informative part. The network is trained on ImageNet, and a large share of ImageNet’s 1,000 categories are dog breeds, something like 120 of them. To get the classification task right, the network had to become very good at telling a Norfolk terrier from a Norwich terrier, so it built a lot of machinery for dog faces. Ask it to amplify whatever it sees and dog-face detectors fire everywhere, because they’re a large part of what it has.
That’s a useful reminder that a network’s view of the world is shaped by its training set and its task, and it doesn’t come from any more neutral place. We tend to talk about “a network trained on ImageNet” as if it learned vision in general. It learned what it needed to win at one particular labeling task, including a lot of expertise about dogs that says more about how the dataset was put together than about the visual world.
The same logic applies to anything I’d train. A face model trained on mostly well-lit, front-facing photos of young people will be very good at well-lit, front-facing photos of young people, and it will believe the world looks like that. DeepDream just makes the belief visible, which is more than most models do.