Radial basis functions, for people who don't care about radial basis functions
I presented a facial animation paper this week. What it's about, explained with a rubber sheet, some pins and a steel plate.
I presented our paper this week, “Improved radial basis function based parameterization for facial expression animation.” Most of the room wasn’t from graphics. The equation slides got polite nods, and the part people actually responded to was when I held up my hands and pretended to stretch a sheet of rubber, so I’ll use the rubber here too.
You have a 3D face and you know where a few dozen points should move for an expression: mouth corners up and out, brows up, jaw down a little. A rig or a face tracker gives you those. The face has thousands of vertices, and you have to decide where all the others go. You want a function that gives the motion of any point, hits the control points exactly, and is smooth in between, without tearing or folding. It’s the same kind of interpolation as guessing the 2 p.m. temperature from the noon and 4 p.m. readings, except in 3D on a surface.
Picture the face as a rubber sheet with a pin at each control point. Pull the pins to their targets and the rubber between them follows. A radial basis function describes one pin’s pull, strongest at the pin and weaker with distance, and it depends on distance only (hence “radial”). A point’s motion is the sum of all the pulls on it:
where are the control points, is the falloff shape, are weights, and is a small linear term. You solve one linear system, sized by the number of control points, to make the function hit every target, and after that each vertex is just a weighted sum. That’s cheap enough for a phone, which was honestly the main reason I picked RBFs.
Choosing is where it gets interesting. A Gaussian is the obvious choice, but you have to pick a width, and every width is wrong somewhere. Too narrow gives you little dents around each pin, and too wide means pulling a mouth corner also moves the ear. The thin-plate spline is the other usual option, and as a mechanical engineer I like where it comes from. Clamp a thin steel plate, push some points to set heights, and it settles into the shape with minimum bending energy. That shape can be written exactly as a sum of radial basis functions with a specific kernel. Bookstein used it in the late 80s to compare skulls between species, so the same math that maps an ape skull onto a human one also maps your neutral face onto your smile, which I find funny.
On faces the trouble is that straight-line distance isn’t distance along the face. Your upper and lower lips are a few millimeters apart in space but on opposite sides of an opening, and they move in opposite directions when you talk. An RBF only knows straight-line distance, so the upper-lip pin pulls the lower lip too, and the eyelids have the same problem. Our early mouths wouldn’t open properly. It looked like the lips were stuck together with honey.
The “improved” in the title is about that: stopping a control point’s influence from leaking across the mouth and the eyes, so each region follows its own points and the surface stays smooth where the face really is continuous. The paper has the math, if you do care about radial basis functions.
Eventually I expect learned models to replace hand-picked kernels like these. A network that has seen enough faces moving would learn that lips separate without anyone telling it. It would still have to output its answer in some representation though, and whatever we pick will have its own version of the honey-lips problem somewhere.