The metaverse needs faces first
Facebook is now Meta and says it's building the metaverse. From someone who spent years on 3D faces: avatars that can't emote are just chat rooms with legs.
On Thursday Facebook renamed itself Meta and said the company’s future is the metaverse, a set of shared virtual spaces you’d enter through VR and AR headsets. The keynote showed people meeting as avatars, playing cards in space, attending concerts together. Zuckerberg says it will take most of a decade.
I’ll leave the business questions to others. My question is narrower, and it’s the one I spent years on: what do people look like in there?
The avatars in the keynote were cartoonish, upper-body only, with no legs, which people have made fun of all week. The legs aren’t the real problem. The faces are. Most of the information in a conversation between people who can see each other is in the face: gaze, micro-expressions, the timing of a smile, a look that says “go on” or “I’m confused.” That’s why video calls, for all their problems, beat voice calls for most conversations. An avatar with a fixed smile and a mouth that flaps to the audio gives you less than a video call does.
Getting this right has two parts, and both are hard.
The first is capture. A headset covers the top half of your face, which is where the eyes and brows are. To animate an avatar faithfully, you need cameras inside the headset looking at your eyes and mouth, and a model that infers the whole face from partial, oddly angled views. Meta’s research lab in Pittsburgh has shown “codec avatars” that do this with impressive results, but the capture setups for building each person’s avatar involve large rigs with many cameras.
The second is rendering and the uncanny valley. I wrote in 2014 about how our 3D faces looked dead because their timing was wrong. Photorealistic faces in VR will run into the same wall, harder, because you’re looking at them from a meter away in stereo. A slightly wrong photoreal face is worse than an honest cartoon. That’s probably why Meta’s current avatars are cartoons.
My guess is that faces will be the gating problem for social VR, more than headsets, bandwidth or content. The technology is coming together: face tracking in headsets, neural rendering like NeRF, generative models like StyleGAN that know what faces look like. Someone will put them together into avatars that carry real expression. When that happens, being present with someone in VR might start to feel like being there. Until then it’ll feel like a game lobby.