[Paper Review] MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
MeshTalk proposes a generic, audio-driven 3D face animation method that disentangles audio-correlated (e.g., lip movements) and audio-uncorrelated (e.g., blinks, eyebrow motions) facial motion using a categorical latent space and a novel cross-modality loss. It achieves state-of-the-art performance in realism and lip-sync accuracy, with 75% of perceptual study participants preferring it over SOTA baselines.
This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their scalability. To improve upon existing models, we propose a generic audio-driven facial animation approach that achieves highly realistic motion synthesis results for the entire face. At the core of our approach is a categorical latent space for facial animation that disentangles audio-correlated and audio-uncorrelated information based on a novel cross-modality loss. Our approach ensures highly accurate lip motion, while also synthesizing plausible animation of the parts of the face that are uncorrelated to the audio signal, such as eye blinks and eye brow motion. We demonstrate that our approach outperforms several baselines and obtains state-of-the-art quality both qualitatively and quantitatively. A perceptual user study demonstrates that our approach is deemed more realistic than the current state-of-the-art in over 75% of cases. We recommend watching the supplemental video before reading the paper: https://github.com/facebookresearch/meshtalk
Motivation & Objective
- Address the limitation of existing audio-driven 3D face animation methods that produce static or uncanny upper face motion due to insufficient modeling of non-audio-related facial dynamics.
- Overcome the one-to-many mapping challenge in audio-driven animation by disentangling audio-correlated and audio-uncorrelated facial motion components.
- Enable generic, identity-agnostic 3D face animation from speech without requiring person-specific training data or high-fidelity motion capture.
- Improve perceptual realism and co-articulation by modeling plausible, diverse upper face motions (e.g., blinks, eyebrow raises) independent of audio input.
- Support practical applications such as re-targeting and mesh dubbing by learning a disentangled, reusable latent representation of facial motion.
Proposed method
- Introduces a categorical latent space that explicitly disentangles audio-correlated (e.g., mouth shape) and audio-uncorrelated (e.g., eye blinks) facial motion components.
- Employs a novel cross-modality loss that encourages accurate reconstruction of upper face motion independent of audio input and precise lip motion conditioned solely on audio.
- Uses an autoregressive sampling strategy over the learned latent space to generate temporally coherent, high-fidelity 3D facial animations from speech.
- Leverages a single neutral 3D face scan and audio input to generate full-face animations without requiring identity-specific fine-tuning.
- Enables re-targeting and mesh dubbing by mapping latent codes from a source identity’s animated mesh to a target identity’s neutral template mesh.
- Trains the model end-to-end using paired audio and 3D mesh sequences, with the loss function promoting disentanglement and modality-specific accuracy.
Experimental results
Research questions
- RQ1Can a disentangled latent space improve the realism of audio-driven 3D facial animation by decoupling audio-dependent and audio-independent facial motion?
- RQ2How can a cross-modality loss be designed to ensure accurate lip motion while preserving plausible, diverse upper face motion such as blinks and eyebrow movements?
- RQ3To what extent can a generic, identity-agnostic model outperform person-specific models in terms of realism and generalization across unseen identities?
- RQ4Can the disentangled latent representation support practical applications like re-targeting and language-agnostic mesh dubbing?
- RQ5Does the proposed method achieve superior perceptual quality compared to state-of-the-art baselines in user studies?
Key findings
- In a perceptual user study, MeshTalk was preferred over the current SOTA method (VOCA) in 75% of cases, with 66.4% of cases strictly preferred and only 42.1% of cases ranked less favorably than ground truth.
- The model achieves highly accurate lip motion that aligns precisely with speech input, as validated by quantitative metrics and visual inspection.
- Upper face motion such as eye blinks and eyebrow raises is naturally varied and plausible, even in the absence of audio cues, due to disentanglement from audio input.
- The disentangled latent space enables successful re-targeting of facial animations from one identity to another while preserving motion quality and identity-agnostic motion patterns.
- Mesh dubbing is successfully demonstrated by re-synthesizing lip shapes for a new language audio while maintaining original upper face motion, proving modality disentanglement enables language-agnostic animation.
- The approach outperforms multiple baselines in both qualitative and quantitative evaluations, establishing a new state-of-the-art in audio-driven 3D face animation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.