[論文レビュー] Everybody's Talkin': Let Me Talk as You Want
本論文は、ソース音声をターゲット表現パラメータに変換してターゲットの人物肖像ビデオをエンドツーエンドで編集し、ニューラルビデオレンダリングネットワークを用いて写真実写レベルの動画を生成するフレームワークを提示する。これにより、個別のアイデンティティ特化ネットワークを用いず、多対多の音声からビデオへの翻訳が可能となる。
We present a method to edit a target portrait footage by taking a sequence of audio as input to synthesize a photo-realistic video. This method is unique because it is highly dynamic. It does not assume a person-specific rendering network yet capable of translating arbitrary source audio into arbitrary video output. Instead of learning a highly heterogeneous and nonlinear mapping from audio to the video directly, we first factorize each target video frame into orthogonal parameter spaces, i.e., expression, geometry, and pose, via monocular 3D face reconstruction. Next, a recurrent network is introduced to translate source audio into expression parameters that are primarily related to the audio content. The audio-translated expression parameters are then used to synthesize a photo-realistic human subject in each video frame, with the movement of the mouth regions precisely mapped to the source audio. The geometry and pose parameters of the target human portrait are retained, therefore preserving the context of the original video footage. Finally, we introduce a novel video rendering network and a dynamic programming method to construct a temporally coherent and photo-realistic video. Extensive experiments demonstrate the superiority of our method over existing approaches. Our method is end-to-end learnable and robust to voice variations in the source audio.
研究の動機と目的
- 任意のソースとターゲットのアイデンティティに跨る肖像動画の自己回帰型音声ベース編集を動機づけ、実現する。
- 動画フレームを幾何学、姿勢、表情へと分離して、頑健な音声からビデオへの翻訳を促進する。
- 音声 ID除去ネットワークを開発し、音声から表情への翻訳をアイデンティティに依存しないものとする。
- 口元ランドマークに基づく口元補完を行うニューラルビデオレンダリングネットワークを提案し、リアリズムと時間的一貫性を確保する。
提案手法
- モノキュラー3D顔再構成で各フレームを解析し、幾何、表情、姿勢(3DMMベースのパラメータ)を取得する。
- 話者に依存しない音声特徴を表情パラメータへ写像する Audio-to-Expression Translation Network(LSTMベース)を用いる。
- 翻訳前に音声特徴から話者アイデンティティを除去する Audio ID-Removing Network を組み込む。
- 口元領域の生成を、口元ランドマークとマスク付きフレームを条件とする顔補完問題として定式化する。
- Unet風アーキテクチャとランドマークヒートマップを用いたニューラルビデオレンダリングネットワークを適用し、再構成、敵対的、知覚的、時間的損失で訓練する。
実験結果
リサーチクエスチョン
- RQ1任意の話者の音声で、個別の訓練なしに任意のターゲットアイデンティティの写真実写風肖像動画を駆動できるか?
- RQ2動画フレームを幾何、姿勢、表情に分解することは、話者を跨ぐ音声から動画への翻訳を改善するか?
- RQ3アイデンティティに依存しない音声表現は、多様な声での口元同期と視覚的リアリズムを向上させるか?
主な発見
- 提案された3Dパラメータベースのミスアラインメント手法は、複数のアイデンティティに跨る単一の生成器で多対多の音声から動画への翻訳を可能にする。
- Audio ID-Removing Networkは口元同期を向上させ、音声特徴におけるアイデンティティ漏えいを低減する。
- ランドマーク案内付きの補完ベース口元生成を共同訓練すると、PSNR/SSIMが競合的で、ベースラインよりリアリズムが向上する。
- 本手法は大きな姿勢変化と音声編集/歌唱をサポートし、未知の話者と言語へ頑健に一般化する。
- 最先端手法と比較して、多くのシナリオでより良い質感ディテールと背景のシームレスなブレンディングを提供する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。