[論文レビュー] Deep Video Portraits
本論文では、写真のようにリアルな顔貌動画の再アニメーションを可能にする、空間時間アーキテクチャを備えた生成的ニューラルネットワーク「Deep Video Portraits」を紹介する。この手法は、ソース・アクトリストからターゲットに、完全な3次元頭部の動き、表情、視線、瞬きを転送することで実現される。生成手法では、パrametricな顔モデルの合成レンダリングを用いた adversarial 訓練により、髪、身体、背景の明示的モデリングを伴わずに、リアルな動画フレームを生成する。
We present a novel approach that enables photo-realistic re-animation of portrait videos using only an input video. In contrast to existing approaches that are restricted to manipulations of facial expressions only, we are the first to transfer the full 3D head position, head rotation, face expression, eye gaze, and eye blinking from a source actor to a portrait video of a target actor. The core of our approach is a generative neural network with a novel space-time architecture. The network takes as input synthetic renderings of a parametric face model, based on which it predicts photo-realistic video frames for a given target actor. The realism in this rendering-to-video transfer is achieved by careful adversarial training, and as a result, we can create modified target videos that mimic the behavior of the synthetically-created input. In order to enable source-to-target video re-animation, we render a synthetic target video with the reconstructed head animation parameters from a source video, and feed it into the trained network -- thus taking full control of the target. With the ability to freely recombine source and target parameters, we are able to demonstrate a large variety of video rewrite applications without explicitly modeling hair, body or background. For instance, we can reenact the full head using interactive user-controlled editing, and realize high-fidelity visual dubbing. To demonstrate the high quality of our output, we conduct an extensive series of experiments and evaluations, where for instance a user study shows that our video edits are hard to detect.
研究の動機と目的
- 顔貌動画における完全な3次元頭部の再アニメーションを可能にすること。これには、頭部の位置、回転、表情、視線、瞬きを含む。
- 髪、身体、背景の明示的モデリングなしに、写真のようにリアルな動画生成を達成すること。
- エンドツーエンド学習により、インタラクティブでユーザー制御可能な動画の再現と高精細なダビングを可能にすること。
- ソース動画の動きパラメータをターゲット動画に転送するが、アイデンティティとリアリズムを保持する手法の開発
提案手法
- パラメトリックな顔モデルの合成レンダリングから、写真のようにリアルな動画フレームを予測するための、新規の空間時間的生成的ニューラルネットワークアーキテクチャを採用する。
- ソース動画から再構築された頭部アニメーションパラメータに基づいて、合成レンダリングを生成し、ソースからターゲットへの再アニメーションを可能にする。
- 視覚的精細度を高めるために、動画生成プロセスにおいて adversarial 訓練を採用する。
- 制御された動きパラメータを備えた合成入力動画上でネットワークを訓練することで、ターゲット動画出力の精密な制御を可能にする。
- アイデンティティ(ターゲット)と動き(ソース)を分離することで、再トレーニングなしに柔軟にパラメータを再結合可能にする。
- ユーザーがソースの動きパラメータを操作し、それをトレーニング済みのネットワークに供給することで、ユーザー制御による編集を支援する。
実験結果
リサーチクエスチョン
- RQ1深層生成モデルは、回転、視線、瞬きを含む完全な3次元頭部の動きを、ソースからターゲットの顔貌動画に写真のようにリアルな結果で転送できるか?
- RQ2背景や身体の明示的モデリングなしに、合成レンダリングに対する adversarial 訓練が、どの程度動画生成における高リアルさを達成できるか?
- RQ3この手法は、インタラクティブでユーザー制御可能な動画の再現およびビジュアルダビングにおいて、どの程度効果的か?
- RQ4知覚的評価において、生成された動画は実際に撮影された動画と明確に区別できるか?
主な発見
- 本手法は、頭部の回転、視線、瞬きを含む完全な3次元頭部の動きを、高い視覚的精細度でソースからターゲットに転送することに成功した。
- ユーザー研究の結果、生成された動画編集は検出が困難であることが示され、強力な知覚的リアリズムを示している。
- 髪、身体、背景の明示的モデリングなしに、高精細なビジュアルダビングとインタラクティブな再現が可能である。
- 合成レンダリングに対する adversarial 訓練により、生成された動画フレームのリアリズムが顕著に向上した。
- ソースとターゲットの入力パラメータを再結合することで、ターゲット動画挙動に対する完全な制御が達成された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。