Skip to main content
QUICK REVIEW

[论文解读] Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos

Panagiotis P. Filntisis, George Retsinas|arXiv (Cornell University)|Jul 22, 2022
Facial Nerve Paralysis Treatment and Research被引用 7
一句话总结

本文提出了一种新颖的视觉语音感知3D面部表情重建方法,从视频中重建3D说话人脸,通过引入'唇读'损失来提升口部动作的感知真实感,而无需文本转录或音频输入。通过利用预训练的唇读模型,最小化重建结果与真实说话人脸之间的感知差异,该方法在口部运动保真度方面优于几何损失和基于关键点的损失,显著提升了3D说话人脸合成的感知自然度。

ABSTRACT

The recent state of the art on monocular 3D face reconstruction from image data has made some impressive advancements, thanks to the advent of Deep Learning. However, it has mostly focused on input coming from a single RGB image, overlooking the following important factors: a) Nowadays, the vast majority of facial image data of interest do not originate from single images but rather from videos, which contain rich dynamic information. b) Furthermore, these videos typically capture individuals in some form of verbal communication (public talks, teleconferences, audiovisual human-computer interactions, interviews, monologues/dialogues in movies, etc). When existing 3D face reconstruction methods are applied in such videos, the artifacts in the reconstruction of the shape and motion of the mouth area are often severe, since they do not match well with the speech audio. To overcome the aforementioned limitations, we present the first method for visual speech-aware perceptual reconstruction of 3D mouth expressions. We do this by proposing a "lipread" loss, which guides the fitting process so that the elicited perception from the 3D reconstructed talking head resembles that of the original video footage. We demonstrate that, interestingly, the lipread loss is better suited for 3D reconstruction of mouth movements compared to traditional landmark losses, and even direct 3D supervision. Furthermore, the devised method does not rely on any text transcriptions or corresponding audio, rendering it ideal for training in unlabeled datasets. We verify the efficiency of our method through exhaustive objective evaluations on three large-scale datasets, as well as subjective evaluation with two web-based user studies.

研究动机与目标

  • 为解决从视频中重建3D人脸时口部动作缺乏感知真实感的问题,特别是说话过程中的口部动作。
  • 克服传统2D关键点监督与3D监督的局限性,后者无法捕捉与语音相关的发音细节。
  • 开发一种方法,提升重建的3D说话人脸在人类语音感知上的感知一致性。
  • 通过仅使用视觉数据进行训练,消除对文本转录或音频的依赖。
  • 验证感知损失在建模与语音相关的面部动态方面优于几何损失和直接3D监督。

提出的方法

  • 该方法引入了一种'唇读'损失,利用预训练的唇读模型,最小化渲染的3D说话人脸与原始视频画面之间的感知距离。
  • 在优化过程中应用唇读损失,以引导3D网格拟合过程,使重建的口部动作在语音感知上与原始视频保持一致。
  • 该方法将唇读损失与基于关键点的损失相结合,以平衡感知真实感与几何精度,避免因过度优化导致的伪影。
  • 它基于先进的3D人脸重建流水线(如DECA或类似方法),并在此基础上引入感知损失,无需3D真实值或音频输入。
  • 该方法在无标签视频数据上端到端进行训练,适用于大规模真实世界应用。
  • 采用相对关键点损失以稳定训练过程,防止因仅使用唇读损失导致的形状失真。

实验结果

研究问题

  • RQ1基于唇读的感知损失能否提升从视频中重建的3D说话人脸口部动作的真实感?
  • RQ2唇读损失是否在重建与语音相关的面部动态方面优于传统的2D关键点监督和直接3D监督?
  • RQ3能否在不依赖文本转录或音频的情况下,实现高质量的3D人脸重建?
  • RQ4唇读损失与关键点损失的结合如何影响感知真实感与几何保真度之间的权衡?
  • RQ5尽管在语音视频上进行训练,该方法在非语音口部动作上的泛化能力如何?

主要发现

  • 客观与主观评估均证实,唇读损失在重建与语音相关的口部动作方面显著优于传统的2D关键点监督和直接3D监督。
  • 在用户研究中,该方法在感知质量方面表现更优,参与者一致偏好该方法生成的3D说话人脸,而非基线方法的结果。
  • 即使没有文本转录或音频输入,该方法仍能生成保留语音感知特征的3D重建结果,证明了纯视觉感知监督的有效性。
  • 唇读损失使口部动作更加清晰自然,尤其在双唇辅音和圆唇元音上表现更佳,这些特征对感知自然度至关重要。
  • 该方法在非语音面部动作上也表现出良好的泛化能力,表明所学习的表征不仅限于语音,还捕捉了通用的表情动态。
  • 尽管感知性能优异,但客观指标如CER和WER仍偏高,主要由于渲染图像与真实图像之间的领域差异,特别是缺少牙齿和舌头细节。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。