Skip to main content
QUICK REVIEW

[論文レビュー] DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion

Johanna Karras, Aleksander Holynski|arXiv (Cornell University)|Apr 12, 2023
Generative Adversarial Networks and Image Synthesis被引用数 4
ひとこと要約

DreamPoseは、Stable Diffusionをポーズおよび画像条件付きで適応させることで、1枚の静止画と人体ポーズのシーケンスからフォトリアリスティックなファッション動画を生成する拡散ベースの手法である。2重CLIP-VAEエンコーダー、5ポーズ条件付き処理、およびUNetとVAEデコーダーの被写体固有の微調整により、時間的整合性、外観忠実度、動きのリアリズムにおいて最先端の結果を達成している。

ABSTRACT

We present DreamPose, a diffusion-based method for generating animated fashion videos from still images. Given an image and a sequence of human body poses, our method synthesizes a video containing both human and fabric motion. To achieve this, we transform a pretrained text-to-image model (Stable Diffusion) into a pose-and-image guided video synthesis model, using a novel fine-tuning strategy, a set of architectural changes to support the added conditioning signals, and techniques to encourage temporal consistency. We fine-tune on a collection of fashion videos from the UBC Fashion dataset. We evaluate our method on a variety of clothing styles and poses, and demonstrate that our method produces state-of-the-art results on fashion video animation.Video results are available on our project page.

研究の動機と目的

  • ファッション写真の普及にもかかわらず、リアルで時間的に整合性のあるファッション動画の不足に取り組む。
  • 単一の入力画像とポーズシーケンスのみを用いて、人間の動きと生地のダイナミクスを正確に制御する高精細な画像アニメーションを実現する。
  • 従来の動画拡散モデルの限界、例えば時間的整合性の低さやテキスト条件付けに依存する点を克服する。
  • 事前学習済みの拡散モデルにポーズと画像条件付けを統合することで、ファッション動画合成における外観忠実度と動きの滑らかさを向上させる。
  • 膨大なデータや複雑なマルチステージネットワークを必要とせずに、被写体固有のアニメーションを実現する。

提案手法

  • U既是ファッションデータセットで事前学習済みのテキストから画像へのStable Diffusionモデルを、新規の2段階戦略で微調整する:まずUBC Fashionデータセットで、次に単一の入力画像での被写体固有の微調整を行う。
  • CLIPテキストエンコーダーを、画像と潜在表現を共同でモデル化・変形するための2重CLIP-VAE画像エンコーダーとアダプターモジュールに置き換えることで、外観忠実度を向上させる。
  • 推定対象フレームを中心とする連続5フレームのポーズシーケンスを、ノイズ除去用のU-Netに条件付けすることで、時間的整合性を強化する。
  • アテンション機構と入力インジェクションパスを変更することで、画像とポーズの両方の条件付けをサポートする新しいアーキテクチャを導入する。
  • 被写体固有の微調整をUNet、アダプターモジュール、VAEデコーダーに対して実施することで、フレーム間でのアイデンティティとテクスチャの詳細を保持する。
  • 各ノイズ除去ステップで、ポーズ表現を入力ノイズに連結するノイズインジェクション機構を採用する。
Figure 2: Architecture Overview. We modify the original Stable Diffusion architecture in order to enable image and pose conditioning. First, we replace the CLIP text encoder with a dual CLIP-VAE image encoder and adapter module (shown in the blue box). The adapter module jointly models and reshapes
Figure 2: Architecture Overview. We modify the original Stable Diffusion architecture in order to enable image and pose conditioning. First, we replace the CLIP text encoder with a dual CLIP-VAE image encoder and adapter module (shown in the blue box). The adapter module jointly models and reshapes

実験結果

リサーチクエスチョン

  • RQ1事前学習済みの画像拡散モデルは、単一の画像とポーズシーケンスから、高品質でフォトリアリスティックなファッション動画を効果的に生成するために適応可能か?
  • RQ2複数の連続ポーズによる条件付けは、単一ポーズ条件付けと比較して、時間的整合性をどのように向上させるか?
  • RQ3UNetおよびVAEデコーダーの被写体固有の微調整は、外観忠実度の向上とフレーム間のちらつきの低減に、どの程度寄与するか?
  • RQ4標準のCLIP画像エンコーダーと比較して、2重CLIP-VAEエンコーダーは、衣類の細部とアイデンティティの保持にどの程度優れているか?
  • RQ5本手法の失敗モードは何か?また、改善されたポーズ推定または追加の監視により、それらを緩和できるか?

主な発見

  • 完全なDreamPoseモデルは、すべての指標で最高のパフォーマンスを達成:L1 (0.019)、SSIM (0.900)、VGG (0.207)、LPIPS (0.056)、アンサンブル変種を上回る。
  • CLIP画像エンコーダーのみを用いたアブレーション(Ours_CLIP)では顕著な外観劣化が見られ、細部の保持には2重CLIP-VAEエンコーダーの必要性が示された。
  • VAEデコーダーの被写体固有の微調整は、シャープネスの向上と過学習の低減に顕著に寄与し、Ours_No-VAE-FTはOurs_CLIPと比較して指標が向上した。
  • 1つのポーズのみを用いた場合(Ours_1-pose)は、定量的スコアは類似しているものの、ちらつきと動きの不安定さが顕著となり、複数フレームのポーズコンテキストの重要性が示された。
  • 1ポーズ変種に時間的スムージングを適用した場合(Ours_smooth)は、性能が著しく低下し、運動アーチファクトは後処理ではなく、アーキテクチャ設計によってより効果的に解消可能であることが示された。
  • 失敗事例には、腕が服に溶け込む、幻覚的な特徴、背面ポーズにおける方向性のずれが含まれ、ポーズ推定および幾何的推論の限界が示唆された。
Figure 3: Qualitative Results. We showcase the results of our method on a variety of input frames and poses. DreamPose is capable of synthesizing photorealistic video frames consistent with a diverse range of patterns, fabric types, person identities, clothing shapes, and viewpoints.
Figure 3: Qualitative Results. We showcase the results of our method on a variety of input frames and poses. DreamPose is capable of synthesizing photorealistic video frames consistent with a diverse range of patterns, fabric types, person identities, clothing shapes, and viewpoints.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。