Vinson·Li

Essay No. 109

Vision Pro scans your face to make you

Apple's headset builds a 3D 'Persona' of your face so you can appear on video calls while wearing it. Ten years after my thesis, 3D face capture ships inside a headset.


Apple announced Vision Pro on Monday. It’s a $3,499 headset, shipping early next year, that Apple insists on calling a spatial computer. Most of the keynote was about windows floating in your living room, eye tracking plus a pinch of the fingers as the input, and watching movies on a huge virtual screen.

The part I’ve been thinking about for three days is a short segment on FaceTime.

When you’re wearing Vision Pro, the people you’re calling can’t see your face, since it’s covered by the headset. Apple’s answer is a “Persona.” During setup, you take the headset off and hold it up to scan your face with its front sensors. It builds a 3D digital version of you using what Apple called an advanced encoder-decoder neural network. During calls, the headset’s internal cameras track your eyes and the sensors track your face and hands, and they drive your Persona in real time. The other person sees a realistic-looking version of you, moving and talking.

There’s also EyeSight: an outward-facing display on the front of the headset that shows a rendering of your eyes to the people in the room with you, so they can tell whether you’re looking at them.

This is exactly the problem I wrote about in 2021, when Facebook became Meta: social VR needs faces, the headset covers the top half of your face, and a slightly wrong photoreal face is worse than an honest cartoon. Apple chose photoreal. The reactions to the Persona footage in the keynote are mixed. The word “uncanny” comes up a lot. It looks like it lives right at the bottom of the valley I wrote about in 2014: very close, and wrong in the motion and the eyes.

Still, it’s striking to see it ship. My master’s thesis in 2014 was about building an animatable 3D face from a scan of a person, so that one template rig could make any face move. The pipeline Apple described is that, done with a neural network instead of a template, with far better sensors, and driven by live tracking from inside the headset. Ten years ago it was a research project. Next year it’s a setup step for a consumer product.

I’d expect Personas to get noticeably better through software updates, since the hard part now is the learned rendering and animation, which improve with data. The eyes will be the last thing to feel right. They always are.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…