[論文レビュー] PhotoApp: Photorealistic Appearance Editing of Head Portraits
PhotoAppは、ライトステージからの限定的な教師ありデータを用いて、StyleGANの潜在空間における変換を学習することで、写真のようにリアルな顔のポートレート編集を実現する新規手法を提示する。本手法は、インタラクティブな速度で、実世界の画像に対して高精細かつ同時に照明の再設定とポーズ編集を可能にし、最先端のリアルさと一般化性能を達成する。定量的指標および視覚的品質の両面で、既存手法を上回る。
Photorealistic editing of portraits is a challenging task as humans are very sensitive to inconsistencies in faces. We present an approach for high-quality intuitive editing of the camera viewpoint and scene illumination in a portrait image. This requires our method to capture and control the full reflectance field of the person in the image. Most editing approaches rely on supervised learning using training data captured with setups such as light and camera stages. Such datasets are expensive to acquire, not readily available and do not capture all the rich variations of in-the-wild portrait images. In addition, most supervised approaches only focus on relighting, and do not allow camera viewpoint editing. Thus, they only capture and control a subset of the reflectance field. Recently, portrait editing has been demonstrated by operating in the generative model space of StyleGAN. While such approaches do not require direct supervision, there is a significant loss of quality when compared to the supervised approaches. In this paper, we present a method which learns from limited supervised training data. The training images only include people in a fixed neutral expression with eyes closed, without much hair or background variations. Each person is captured under 150 one-light-at-a-time conditions and under 8 camera poses. Instead of training directly in the image space, we design a supervised problem which learns transformations in the latent space of StyleGAN. This combines the best of supervised learning and generative adversarial modeling. We show that the StyleGAN prior allows for generalisation to different expressions, hairstyles and backgrounds. This produces high-quality photorealistic results for in-the-wild images and significantly outperforms existing methods. Our approach can edit the illumination and pose simultaneously, and runs at interactive rates.
研究の動機と目的
- シーンの照明とカメラの視点を同時に変更することで、高品質で写真のようにリアルな顔のポートレート編集を可能にすること。
- 従来の教師あり手法が膨大なデータを必要とし、照明やポーズのいずれかの側面のみを一度に編集できることの制限を克服すること。
- 訓練データにない顔の表情、ヘアスタイル、背景の多様性に対しても一般化できること。
- StyleGANの潜在空間における最適化を通じて、教師あり学習と生成モデルの長所を統合すること。
- 高い視覚的忠実性と一貫性を維持しながら、インタラクティブな推論速度を達成すること。
提案手法
- 本手法は、ピクセル空間ではなく、StyleGANの潜在空間における潜在コード変換を予測するためのニューラルネットワークを学習する。
- 中立的な表情と閉じた目を有する被験者150名分の、1人あたり150通りのライトステージ条件下(1つのライトのみ)と、8つのカメラポーズを含む、制御された小規模なライトステージデータセットを用いる。
- ネットワークは、元の画像の潜在コードを、ターゲットの照明(環境マップを介して)と視点を反映する新しいコードにマッピングする。
- 照明変更の有無を区別するためのバイナリ条件を導入し、入力の照明を保持したままポーズ編集を独立して行えるようにする。
- StyleGANの強力な事前知識を活用することで、訓練データにない未確認の被験者、表情、背景に対しても一般化可能である。
- 本手法はリアルタイムで動作し、顔のポートレート画像のインタラクティブ編集を可能にする。
実験結果
リサーチクエスチョン
- RQ1小規模な教師ありライトステージデータセットを用いて、多様な実世界の顔のポートレートに一般化可能なモデルを学習できるか?
- RQ2StyleGANにおける潜在空間編集は、直接的なピクセル空間学習に比べて、より高い写真的リアリズムを達成できるか?
- RQ3高精細かつ一貫性のある高品質な、同時に照明の再設定とポーズ編集が可能か?
- RQ4顔の識別性および髪の毛や目などの細部が、大規模な照明変更やポーズ変更に対しても保持されるか?
- RQ5微調整なしで、未確認の表情、ヘアスタイル、背景に対しても一般化できるか?
主な発見
- 全テストセット(150名分)において、Si-MSEが0.0020、SSIMが0.9199を達成し、先行研究を両指標で上回った。
- 訓練に3名分の被験者しか使用しなくても、Si-MSEが0.0020、SSIMが0.9191を達成し、強力な少データ一般化性能を示した。
- 複数のテスト条件において、Tewariら(2020b)およびAbdalら(2020)を著しく上回った。
- 定性的な結果では、新しい照明とポーズ下でも、正確な自己影、体内部透過散乱、鏡面反射光を再現し、高い写真的リアリズムを実現した。
- 訓練中に見られなかった多様な表情、ヘアスタイル、背景を持つ実世界の画像に対しても、本手法は成功裏に一般化した。
- 本手法はインタラクティブな編集速度を達成しており、バーチャルプロダクションやVRにおけるリアルタイム応用に適している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。