Skip to main content
QUICK REVIEW

[논문 리뷰] FLAME-in-NeRF : Neural control of Radiance Fields for Free View Face Animation

ShahRukh Athar, Zhixin Shu|arXiv (Cornell University)|2021. 08. 10.
Generative Adversarial Networks and Image Synthesis인용 수 5
한 줄 요약

FLAME-in-NeRF는 모바일 기기로 촬영한 영상만으로도 얼굴 표정의 명시적이고 분리된 제어 및 포트레이트 영상의 새로운 시점 합성 기능을 가능하게 하는 새로운 신경 렌디언스 필드(NeRF) 방법을 제안한다. 3D 모러포블 모델(3DMM)의 표정 파라미터에 조건을 부여하고 3DMM 피팅에서 유도된 공간적 사전 지식을 적용하여, 다양한 시점에서 정확한 표정 전달을 포함한 고해상도 재생동화를 달성하며, 기존의 NeRF 기반 방법들보다 재구성 및 재생동화 품질에서 뛰어난 성능을 보인다.

ABSTRACT

This paper presents a neural rendering method for controllable portrait video synthesis. Recent advances in volumetric neural rendering, such as neural radiance fields (NeRF), has enabled the photorealistic novel view synthesis of static scenes with impressive results. However, modeling dynamic and controllable objects as part of a scene with such scene representations is still challenging. In this work, we design a system that enables both novel view synthesis for portrait video, including the human subject and the scene background, and explicit control of the facial expressions through a low-dimensional expression representation. We leverage the expression space of a 3D morphable face model (3DMM) to represent the distribution of human facial expressions, and use it to condition the NeRF volumetric function. Furthermore, we impose a spatial prior brought by 3DMM fitting to guide the network to learn disentangled control for scene appearance and facial actions. We demonstrate the effectiveness of our method on free view synthesis of portrait videos with expression controls. To train a scene, our method only requires a short video of a subject captured by a mobile device.

연구 동기 및 목표

  • 명시적 얼굴 표정 제어 기능을 갖춘 포트레이트 영상의 제어 가능하고 사실적인 새로운 시점 합성을 가능하게 하는 것.
  • 동적인 인간 얼굴을 애니메이션할 때 신경 렌디언스 필드 내에서 표정과 외관이 뒤섞이는 문제를 해결하는 것.
  • 특수 하드웨어 없이도 짧은 모바일 기기로 촬영한 영상에서 학습 가능한 시스템을 개발하는 것.
  • 3DMM 기반의 공간적 사전 지식을 활용해 얼굴 표정과 장면 외관을 분리하여, 표정 변화가 영향을 주는 것은 오직 얼굴 관련 기하학적 구조에 국한되도록 보장하는 것.

제안 방법

  • 신경 렌디언스 필드**(NeRF)**를 3DMM 표정 파라미터(γ)에 조건화하여 명시적 얼굴 표정 제어를 가능하게 하는 것.
  • 3DMM 피팅에서 유도된 공간적 레이 샘플링 사전 지식을 통합하여 네트워크가 표정과 장면 외관을 분리하도록 제약을 가하는 것.
  • 학습 안정성을 확보하기 위해 최적화된 변형 코드를 갖춘 캐논리컬 프레임을 사용해 동적 프레임을 캐논리컬 기하학에 정렬하는 것.
  • 픽셀 수준 재구성 손실, 지각 손실, 그리고 표정 및 외관 코드에 대한 정규화 손실로 구성된 다중 구성 요소 손실을 사용해 NeRF를 훈련하는 것.
  • 드라이빙 영상에서 표정 파라미터를 추출하기 위해 DECA를 활용하며, 이를 FLAME-in-NeRF의 재생동화를 유도하는 입력으로 사용하는 것.
  • 계층적 볼륨 샘플링과 볼륨 렌더링을 적용하여, 훈련된 NeRF로부터 표정 및 시점 제어 기능을 갖춘 새로운 시점 렌더링을 수행하는 것.
Figure 1 : FLAME- in -NeRF. Our method, FLAME- in -NeRF, models portrait videos (left) using an expression conditioned neural radiance field with a spatial prior (middle). Once trained, FLAME- in -NeRF can reanimate the subject and the scene present in the portrait video with arbitrary facial expres
Figure 1 : FLAME- in -NeRF. Our method, FLAME- in -NeRF, models portrait videos (left) using an expression conditioned neural radiance field with a spatial prior (middle). Once trained, FLAME- in -NeRF can reanimate the subject and the scene present in the portrait video with arbitrary facial expres

실험 결과

연구 질문

  • RQ1신경 렌디언스 필드가 얼굴 표정 파라미터에 명시적으로 조건화될 수 있는가? 이를 통해 제어 가능한 포트레이트 영상 재생동화가 가능해지는가?
  • RQ23DMM 기반의 공간적 사전 지식이 NeRF 내에서 얼굴 표정과 장면 외관 간의 분리도를 어떻게 향상시키는가?
  • RQ3특수 장비 없이도 짧은 모바일 기기로 촬영한 포트레이트 영상에서 고해상도의 새로운 시점 합성과 표정 제어를 달성할 수 있는가?
  • RQ4표정-외관 뒤섞임 현상이 기존 NeRF 기반 방법에서 재생동화 품질을 얼마나 악화시키며, 이를 효과적으로 완화할 수 있는가?
  • RQ5제안된 방법이 다양한 시점에서 표정 기반 재생동화 중에도 미세한 얼굴 특징(예: 머리카락, 안경 등)을 얼마나 잘 유지하는가?

주요 결과

  • FLAME-in-NeRF는 재구성 및 재생동화 품질에서 Nerfies를 능가하며, 검증 이미지에서 더 낮은 MSE와 더 높은 PSNR를 기록한다.
  • 이 방법은 다양한 시점 방향으로도 고해상도의 표정 전달을 유지하지만, Nerfies는 표정-외관 뒤섞임으로 인해 정확한 표정을 포착하지 못한다.
  • 재생동화 결과에서 FLAME-in-NeRF는 카메라 시점에 관계없이 드라이빙 표정을 정확히 재현하는 반면, Nerfies는 시점에 따라 약간의 변화만 보이다가 표정 변화를 제대로 반영하지 못한다.
  • 3DMM 피팅에서 유도된 공간적 사전 지식은 분리도를 크게 향상시켜, 표정 변화가 장면 외관에 영향을 주는 것을 방지한다.
  • 단지 짧은 모바일 기기로 촬영한 영상만으로도 높은 품질의 결과를 달성하여 소비자 수준의 데이터 촬영 가능성에 대한 타당성을 입증한다.
  • 표정 변화가 크더라도 머리카락, 수염, 이가, 안경과 같은 미세한 얼굴 특징을 재생동화 과정에서도 고해상도로 유지한다.
Figure 2 : Overview of training FLAME- in -NeRF. First, we use DECA [ 10 ] and landmark fitting [ 14 ] to extract per-frame camera, shape, and expression parameters. Next, these parameters are used to render a silhouette of the FLAME model geometry. This silhouette is used to provide a spatial prior
Figure 2 : Overview of training FLAME- in -NeRF. First, we use DECA [ 10 ] and landmark fitting [ 14 ] to extract per-frame camera, shape, and expression parameters. Next, these parameters are used to render a silhouette of the FLAME model geometry. This silhouette is used to provide a spatial prior

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.