Skip to main content
QUICK REVIEW

[논문 리뷰] Everybody's Talkin': Let Me Talk as You Want

Linsen Song, Wayne Wu|arXiv (Cornell University)|2020. 01. 15.
Generative Adversarial Networks and Image Synthesis참고 문헌 59인용 수 35
한 줄 요약

이 논문은 소스 오디오를 목표 표현 매개변수로 변환한 다음 신경 비디오 렌더링 네트워크로 포토리얼리스틱 비디오를 생성하여 엔드투엔드 프레임워크를 제시하며, 정체성 특정 네트워크 없이 다대다 오디오-비디오 변환을 가능하게 한다.

ABSTRACT

We present a method to edit a target portrait footage by taking a sequence of audio as input to synthesize a photo-realistic video. This method is unique because it is highly dynamic. It does not assume a person-specific rendering network yet capable of translating arbitrary source audio into arbitrary video output. Instead of learning a highly heterogeneous and nonlinear mapping from audio to the video directly, we first factorize each target video frame into orthogonal parameter spaces, i.e., expression, geometry, and pose, via monocular 3D face reconstruction. Next, a recurrent network is introduced to translate source audio into expression parameters that are primarily related to the audio content. The audio-translated expression parameters are then used to synthesize a photo-realistic human subject in each video frame, with the movement of the mouth regions precisely mapped to the source audio. The geometry and pose parameters of the target human portrait are retained, therefore preserving the context of the original video footage. Finally, we introduce a novel video rendering network and a dynamic programming method to construct a temporally coherent and photo-realistic video. Extensive experiments demonstrate the superiority of our method over existing approaches. Our method is end-to-end learnable and robust to voice variations in the source audio.

연구 동기 및 목표

  • 임의의 소스 및 대상 정체성에 걸친 포트레이트 비디오의 자기회귀 기반 오디오 편집을 동기 부여하고 가능하게 한다.
  • 비디오 프레임을 기하학, 자세, 표정으로 분리하여 견고한 오디오-비디오 변환을 촉진한다.
  • 오디오-표정 변환을 정체성 불문으로 만드는 Audio ID-Removing Network를 개발한다.
  • 입 주변의 현실감과 시간적 일관성을 보장하기 위해 입 랜드마크로 안내되는 입 분절 영역 완성용 신경 비디오 렌더링 네트워크를 제안한다

제안 방법

  • 각 프레임을 단안 3D 얼굴 재구성으로 해석하여 기하, 표정, 자세(3DMM 기반 매개변수)를 얻는다.
  • 발화자-배제 오디오 특징을 표정 매개변수로 매핑하는 Audio-to-Expression Translation Network(LSTM 기반)를 사용한다.
  • 번역 전에 오디오 특징에서 화자 정체성을 제거하는 Audio ID-Removing Network를 도입한다.
  • 입 주변 영역 생성을 입 랜드마크와 마스킹된 프레임에 조건화된 얼굴 보정 문제로 형식화한다.
  • 재구성, 적대적, 인지적, 시간적 손실로 학습되는 Unet 유사 구조와 랜드마크 히트맵을 갖춘 신경 비디오 렌더링 네트워크를 적용한다

실험 결과

연구 질문

  • RQ1임의의 화자 오디오가 개인별 학습 없이 임의의 대상 정체성의 포토리얼리스틱 포트레이트 비디오를 구동할 수 있는가?
  • RQ2비디오 프레임을 기하, 자세, 표정으로 분해하는 것이 화자 간 오디오-비디오 변환을 개선하는가?
  • RQ3정체성에 구애받지 않는 오디오 표현이 다양한 음성에서 입 모양 동기화와 시각적 리얼리즘을 개선하는가?

주요 결과

  • 제안된 3D 매개변수 기반 정합 벗어남 접근법은 다수의 정체성에 대해 단일 생성기로 다대다 오디오-비디오 변환을 가능하게 한다.
  • Audio ID-Removing Network는 입 모양 동기화를 향상시키고 오디오 특징의 정체성 누출을 줄인다.
  • 함께 학습된 입 주변 생성 기반으로 랜드마크 가이던스를 적용하면 기준 대비 경쟁적인 PSNR/SSIM과 향상된 리얼리즘을 얻을 수 있다.
  • 이 방법은 큰 포즈 변형과 오디오 편집/노래를 지원하며, 보지 못한 화자 및 언어에 대한 견고한 일반화를 보여준다.
  • 최신 방법과 비교하여 다양한 상황에서 질감 디테일과 배경의 매끄러운 혼합이 더 우수하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.