Skip to main content
QUICK REVIEW

[논문 리뷰] Generalizable One-shot Neural Head Avatar

Xueting Li, Shalini De Mello|arXiv (Cornell University)|2023. 06. 14.
Face recognition and analysis인용 수 7
한 줄 요약

이 논문은 테스트 시 최적화 없이 단일 포트레이트 이미지에서 3D 헤드 애벌러를 재구성하고 애니메이션하는 one-shot 신경 헤드 애벌러 방법을 제안한다. 군중 기하학, 세부 외관, 표정을 분리한 삼중 평면 브랜치를 사용하여, 예측되지 않은 신원과 자세에 대해 최신 기술 수준의 성능을 달성한다.

ABSTRACT

We present a method that reconstructs and animates a 3D head avatar from a single-view portrait image. Existing methods either involve time-consuming optimization for a specific person with multiple images, or they struggle to synthesize intricate appearance details beyond the facial region. To address these limitations, we propose a framework that not only generalizes to unseen identities based on a single-view image without requiring person-specific optimization, but also captures characteristic details within and beyond the face area (e.g. hairstyle, accessories, etc.). At the core of our method are three branches that produce three tri-planes representing the coarse 3D geometry, detailed appearance of a source image, as well as the expression of a target image. By applying volumetric rendering to the combination of the three tri-planes followed by a super-resolution module, our method yields a high fidelity image of the desired identity, expression and pose. Once trained, our model enables efficient 3D head avatar reconstruction and animation via a single forward pass through a network. Experiments show that the proposed approach generalizes well to unseen validation datasets, surpassing SOTA baseline methods by a large margin on head avatar reconstruction and animation.

연구 동기 및 목표

  • 테스트 시 최적화 없이 단일 포트레이트 이미지에서 고해상도, 일반화 가능한 3D 헤드 애벌러를 생성하는 도전 과제를 해결한다.
  • 기존 방법의 한계를 극복하여 얼굴 외부 영역(예: 헤어스타일, 안경 등)의 정교한 외관 세부 정보를 유지하지 못하거나 예측되지 않은 신원에서 어려움을 겪는 문제를 해결한다.
  • 추론 시 시간이 오래 걸리는 최적화 없이도 효율적인 단일 전방 전파 추론을 가능하게 하여 3D 헤드 애벌러 재구성과 애니메이션을 구현한다.
  • 거시 기하학, 세부 외관, 얼굴 표정을 별도로 분리하여 모델링하여 정확도와 일반화 능력을 향상시킨다.

제안 방법

  • 중립 표정과 정면 자세를 가진 캐논리컬 브랜치를 사용하여 삼중 평면 표현을 통해 거시 기하학을 재구성한다.
  • 캐논리컬 브랜치에서의 깊이 정보를 활용해 소스 이미지의 픽셀 값을 3D 캐논리컬 공간에 매핑하는 외관 브랜치를 도입하여 세부 정보를 유지한다.
  • 목표 표정과 소스 신원을 가진 정면 3DMM 렌더링을 입력으로 받아 얼굴 표정을 수정하는 삼중 평면을 생성하는 표정 브랜치를 개발한다.
  • 세 개의 삼중 평면을 덧셈으로 조합하고 체적 렌더링을 적용한 후 슈퍼해상도 모듈을 통해 고해상도 출력 이미지를 생성한다.
  • 3DMM 사전 지식을 활용해 형태와 표정의 분리 구조를 구현하며, 기하학 및 표정 파라미터에 대해 선형 기저 분해를 사용한다.
  • 다양한 포트레이트 이미지로 모델을 훈련시어 추론 시 예측되지 않은 신원과 자세에 대해 제로샷 일반화를 가능하게 한다.
Figure 1: Overview. The proposed method contains four main modules: a canonical branch that reconstructs the coarse geometry and texture of a portrait with a neutral expression (Sec. 3.1 ), an appearance branch that captures fine-grained person-specific details (Sec. 3.2 ), an expression branch that
Figure 1: Overview. The proposed method contains four main modules: a canonical branch that reconstructs the coarse geometry and texture of a portrait with a neutral expression (Sec. 3.1 ), an appearance branch that captures fine-grained person-specific details (Sec. 3.2 ), an expression branch that

실험 결과

연구 질문

  • RQ1단일 포트레이트 이미지로 훈련된 신경 헤드 애벌러 모델이 테스트 시 최적화 없이도 예측되지 않은 신원에 일반화될 수 있는가?
  • RQ2단일 이미지 기반 방법이 얼굴 영역 외부의 정교한 외관 세부 정보(예: 헤어스타일, 액세서리 등)를 어느 정도 유지할 수 있는가?
  • RQ3입력 이미지의 가시성이 제한되었을 때, 모델이 가려진 영역, 이가, 동공 등의 합리적인 기하학을 얼마나 잘 환기할 수 있는가?
  • RQ4기하학, 외관, 표정을 별도의 삼중 평면 브랜치로 분리함으로써 종합적 모델링 방법에 비해 재구성 정확도와 일반화 능력이 향상되는가?

주요 결과

  • 제안된 방법은 헤드 애벌러 재구성 및 애니메이션에서 최신 기술 수준의 베이스라인을 뛰어넘어, 예측되지 않은 검증 데이터셋에서 뛰어난 성능을 달성한다.
  • 모델은 추론 시 어떤 최적화도 필요 없이 예측되지 않은 신원과 자세에 효과적으로 일반화되며, 단일 전방 전파 생성을 통해 효율적인 구현이 가능하다.
  • CelebA 실험에서 ROME 및 HeadNeRF와 비교해 강력한 성능을 보였으며, 테스트 시 적응 없이도 FID 및 LPIPS 지표에서 정량적 향상을 보였다.
  • 모델은 다양한 테스트 신원에서 헤어스타일, 안경, 귀걸이 등 신원 고유의 세부 정보를 성공적으로 유지한다.
  • 한계점으로는 표정 브랜치에서 환기로 인해 이가와 동공의 재현이 일관되지 않으며, 입력의 가시성과 관계없이 항상 이러한 기능을 생성한다는 점이다.
  • 제거 실험을 통해 삼중 평면을 통한 기하학, 외관, 표정의 분리가 종합적 모델링 방법에 비해 정확도와 일반화 능력을 크게 향상시킨다는 점을 확인했다.
Figure 2: Visualization of the contribution of each branch. (a) Source image. (b) Target image. (c) Rendering of the canonical tri-plane. (d) Rendering of the combination of the canonical and expression tri-planes. (e) Rendering of the combination of all three tri-planes.
Figure 2: Visualization of the contribution of each branch. (a) Source image. (b) Target image. (c) Rendering of the canonical tri-plane. (d) Rendering of the combination of the canonical and expression tri-planes. (e) Rendering of the combination of all three tri-planes.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.