Skip to main content
QUICK REVIEW

[論文レビュー] Pivotal Tuning for Latent-based Editing of Real Images

Daniel Roich, Ron Mokady|arXiv (Cornell University)|Jun 10, 2021
Generative Adversarial Networks and Image Synthesis被引用数 4
ひとこと要約

本稿では、ピボット潜在コードの周囲で生成器を微調整することにより、StyleGANにおける高精度でアイデンティティを保持する顔画像編集を可能にするPivotal Tuning Inversion(PTI)を提案する。生成器をドメイン外の画像をよりよく表現できるように局所的に調整することで、歪みと編集可能性のトレードオフを解消し、実世界の顔画像において定量的指標および定性的な編集の両面で最先端の結果を達成した。

ABSTRACT

Recently, a surge of advanced facial editing techniques have been proposed that leverage the generative power of a pre-trained StyleGAN. To successfully edit an image this way, one must first project (or invert) the image into the pre-trained generator's domain. As it turns out, however, StyleGAN's latent space induces an inherent tradeoff between distortion and editability, i.e. between maintaining the original appearance and convincingly altering some of its attributes. Practically, this means it is still challenging to apply ID-preserving facial latent-space editing to faces which are out of the generator's domain. In this paper, we present an approach to bridge this gap. Our technique slightly alters the generator, so that an out-of-domain image is faithfully mapped into an in-domain latent code. The key idea is pivotal tuning - a brief training process that preserves the editing quality of an in-domain latent region, while changing its portrayed identity and appearance. In Pivotal Tuning Inversion (PTI), an initial inverted latent code serves as a pivot, around which the generator is fined-tuned. At the same time, a regularization term keeps nearby identities intact, to locally contain the effect. This surgical training process ends up altering appearance features that represent mostly identity, without affecting editing capabilities. We validate our technique through inversion and editing metrics, and show preferable scores to state-of-the-art methods. We further qualitatively demonstrate our technique by applying advanced edits (such as pose, age, or expression) to numerous images of well-known and recognizable identities. Finally, we demonstrate resilience to harder cases, including heavy make-up, elaborate hairstyles and/or headwear, which otherwise could not have been successfully inverted and edited by state-of-the-art methods.

研究の動機と目的

  • ドメイン外の顔画像において、アイデンティティの喪失やアーティファクトが生じる、顔画像の潜在空間ベース編集における歪みと編集可能性のトレードオフを解消すること。
  • 元のStyleGAN生成器の分布外にある実画像に対して、高品質でアイデンティティを保持する編集を可能にすること。
  • 生成器の編集能力を維持しながら再構成精度を向上させる、パーソナライズ可能で効率的な微調整手法を開発すること。
  • 濃いメイク、複雑なヘアスタイル、帽子などの困難なケースにおける堅牢性を示すこと。
  • メディア制作やパーソナライズド画像操作の分野における実用的フレームワークを提供すること。

提案手法

  • 本手法は、既存の手法を用いて実画像を生成器の潜在空間にインバースすることで、W+空間内の初期潜在コードを生成することから始まる。
  • W空間からのピボット潜在コードが選択され、生成器の微調整のためのアンカーとして機能する。
  • 生成器は、局所的な潜在空間領域における編集可能性を保持すると同時に歪みを最小限に抑えるよう、短時間の最適化プロセス(ピボットチューニング)によって微調整される。
  • 局所的なアイデンティティ保持を促進する正則化項が導入され、近接する潜在コードがアイデンティティを維持するよう制約され、全体の分布シフトを防ぐ。
  • 生成器の層は、L2とLPIPS再構成損失に加え、アイデンティティ正則化を組み合わせた損失関数を最適化することでチューニングされる。
  • 最終的な生成器は入力画像にパーソナライズされ、高精度で歪みの少ない編集が可能になる。
Figure 2: An illustration of the PTI method. StyleGAN’s latent space is portrayed in two dimensions (see Tov et al. [ 38 ] ), where the warmer colors indicate higher densities of $W$ , i.e. regions of higher editability. On the left, we illustrate the generated samples before pivotal tuning. We can
Figure 2: An illustration of the PTI method. StyleGAN’s latent space is portrayed in two dimensions (see Tov et al. [ 38 ] ), where the warmer colors indicate higher densities of $W$ , i.e. regions of higher editability. On the left, we illustrate the generated samples before pivotal tuning. We can

実験結果

リサーチクエスチョン

  • RQ1ドメイン外の実画像の潜在空間ベース編集において、編集可能性を損なわず、再構成歪みを効果的に低減できるか?
  • RQ2局所的かつ生成器固有の微調整プロセスは、StyleGANの元の編集能力を保持しつつ、新たなアイデンティティを適応可能にできるか?
  • RQ3ピボットチューニングは、標準的なインバージョン手法と比較して、困難な顔画像におけるアイデンティティ保持と視覚的品質で優れているか?
  • RQ4ピボットチューニングは、複数のアイデンティティに対して歪みが少なく、一貫性のある性能を示せるか?
  • RQ5ピボット潜在コードは初期化に頼らずに安定しており、チューニング中に固定可能で計算コストを削減できるか?

主な発見

  • PTIは、チューニング前後におけるインバースコードの比較において、LPIPS(0.015±5e−6)およびMSE(0.0012±1e−6)のスコアで他の手法を顕著に上回り、歪みが最小限であることを示している。
  • 本手法は、濃いメイク、複雑なヘアスタイル、帽子を装着したケースなど、従来の手法がインバージョンアーティファクトのため失敗する状況においても、高品質な編集を可能にしている。
  • アブレーションスタディの結果、ピボットコードを平均値やランダムな潜在コードに置き換えると歪みが顕著に増加し、ピボット選択の重要性が裏付けられた。
  • PTIは強力な編集可能性を維持しており、セリーナ・ウィリアムズやロバート・ダウニー・ジュニアのような顔が識別可能な人物に対しても、笑顔、ポーズ、年齢、表情などの高度な編集が成功している。
  • 全二段階プロセスは、1枚のRTX 2080で3分未塔で実行可能であり、ピボットチューニングは2分未塔で完了するため、実世界での利用に実用的である。
  • ピボットコードの初期化に頼らずに安定しており、最適化による性能向上がわずかであるため、計算コストを削減する目的で固定可能である。
Figure 4: Reconstruction of out-of-domain samples. Our method (right) reconstructs out-of-domain visual details (left), such as face paintings or hands, significantly better than state-of-the-art methods (middle).
Figure 4: Reconstruction of out-of-domain samples. Our method (right) reconstructs out-of-domain visual details (left), such as face paintings or hands, significantly better than state-of-the-art methods (middle).

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。