A video is not a stack of photos
Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.
Two video papers from this month make more sense read together.
Karpathy et al. at CVPR trained CNNs on about a million YouTube sports videos in 487 classes and compared ways of giving the network time: a single frame, early fusion, late fusion, and “slow fusion,” which merges frames gradually through the layers. Slow fusion won by about a point and a half over the single-frame network. You can tell swimming from bowling in a still, so maybe that’s expected for sports, but it means the network mostly learned appearance.
Simonyan and Zisserman’s two-stream paper from Oxford went up on arXiv a couple of weeks ago. One network looks at a single RGB frame. The other one gets optical flow instead of pixels, meaning the horizontal and vertical displacement of every pixel between consecutive frames, stacked over 10 frames for 20 input channels. They average the two outputs. It gets 88% on UCF-101, roughly level with the best hand-engineered trajectory features, and the flow network by itself does better than the RGB network by itself, which I did not expect. It never sees a color or a texture.
If you haven’t used optical flow: you assume a small patch keeps its brightness while it moves, and that gives you, per pixel,
with spatial gradients , , temporal change , and the velocity as the unknowns. That’s one equation for two unknowns, so a single pixel can’t be solved (the aperture problem, also the reason barber poles look like they’re going up). Lucas-Kanade assumes a small neighborhood moves together, Horn-Schunck assumes the whole field is smooth, and newer methods mostly deal with edges and occlusions.
We did something close to this in fluids lab in undergrad, particle image velocimetry. Seed the water with particles, fire two laser pulses a few milliseconds apart, and cross-correlate small windows between the two images to get a velocity field. Optical flow has to do that without anyone putting nice trackable particles in the scene, so it’s much harder.
So more frames didn’t get the network to learn motion, but computing motion and handing it over worked. The information was in the frames already and the network found appearance an easier route.
This makes me doubt the usual way of handling video now, where you sample a few frames, run an image model on each and average. For “what is this video about” that’s fine. For what happened in it, like whether she hesitated before answering, it’s weak. I work on facial animation and the difference between a real smile and a polite one is mostly timing: how fast the mouth corners rise and whether the eyes come in a bit later. Frozen at the peak they can look almost the same.
I doubt hand-computed flow lasts, though. It’s expensive, and it’s the same kind of hand-designed intermediate step as the landmarks I complained about last month. My guess is that in a few years networks learn motion directly, probably with convolutions that span time as well as space, trained on far more video than has been labeled so far. What should carry over from this paper is making motion an explicit input to the architecture.
YouTube said last year that 100 hours of video are uploaded every minute, and Vine and Instagram now have everyone shooting six and fifteen second clips. Most of it is understood through its title, thumbnail and click data and not much else. Better video understanding would help search and recommendations a lot, and I suspect nobody gets to generating good video without it either, because generating motion needs some model of how things move.