Skip to main content
QUICK REVIEW

[論文レビュー] Understanding the Impact of Negative Prompts: When and How Do They Take Effect?

Yuanhao Ban, Ruochen Wang|arXiv (Cornell University)|Jun 5, 2024
Resilience and Mental HealthPsychology被引用数 3
ひとこと要約

この論文は、テキストから画像への拡散モデルにおけるネガティブプロンプトの最初の包括的分析を提供し、それらが潜在空間内で顕著な遅延を伴い、中和メカニズムを通じて作用することを明らかにした。著者らは、逆拡散プロセスの中点にネガティブプロンプトを適用することで、モデルの微調整なしに効果的かつ高精細なオブジェクトインpaintingを実現し、除去成功率と画像類似度をそれぞれ最大83%および82.6%向上させた。

ABSTRACT

The concept of negative prompts, emerging from conditional generation models like Stable Diffusion, allows users to specify what to exclude from the generated images.%, demonstrating significant practical efficacy. Despite the widespread use of negative prompts, their intrinsic mechanisms remain largely unexplored. This paper presents the first comprehensive study to uncover how and when negative prompts take effect. Our extensive empirical analysis identifies two primary behaviors of negative prompts. Delayed Effect: The impact of negative prompts is observed after positive prompts render corresponding content. Deletion Through Neutralization: Negative prompts delete concepts from the generated image through a mutual cancellation effect in latent space with positive prompts. These insights reveal significant potential real-world applications; for example, we demonstrate that negative prompts can facilitate object inpainting with minimal alterations to the background via a simple adaptive algorithm. We believe our findings will offer valuable insights for the community in capitalizing on the potential of negative prompts.

研究の動機と目的

  • テキストから画像への拡散モデルにおけるネガティブプロンプトの影響タイミングとメカニズムを理解すること。
  • ネガティブプロンプトが生成に与える影響について、経験的使用を超えた体系的分析の欠如を是正すること。
  • 不要なオブジェクトを除去しながら画像の整合性を保つことができる制御可能なインpainting手法を開発すること。
  • 構造的歪みを避けるために、ネガティブプロンプトを適用する最適タイミングを同定すること。

提案手法

  • 著者らは、拡散ステップごとのクロスアテンションマップを分析し、ネガティブプロンプトが生成に影響し始める臨界ステップを同定した。
  • オブジェクト削除タスク中の推定ノイズパターンを検査することで、潜在空間ダイナミクスを調査した。
  • ネガティブプロンプトが正のプロンプトがコンテンツをレンダリングした後にのみ作用する「遅延効果」と、差し引きノイズ相互作用による「中和」メカニズムを同定した。
  • 逆拡散プロセスの中点にネガティブプロンプトを適用する、新しいインpainting戦略を提案した。この戦略により、早期干渉を回避した。
  • モデルの再トレーニングや推論時変更を一切使用せず、プロンプトスケジューリングに依存した。
  • 評価は、複数のデータセット(COCO、CC、MSVD、Places、Vatex、Nocaps)を用い、GPT-4Vおよび人間のアノテーターによる評価を実施。除去成功率、類似度、相対的パフォーマンスを測定した。
Figure 1 : Illustration on when the negative prompts attend to the "right" place. For example, we consider the face of the person as the "right place" for the "glasses" token. Every row represents an independent diffusion process where the first and the third rows show the tokens in the positive pro
Figure 1 : Illustration on when the negative prompts attend to the "right" place. For example, we consider the face of the person as the "right place" for the "glasses" token. Every row represents an independent diffusion process where the first and the third rows show the tokens in the positive pro

実験結果

リサーチクエスチョン

  • RQ1ネガティブプロンプトは拡散プロセスのどの段階で影響を始め、正のプロンプトと比較してそのタイミングはどのように異なるか?
  • RQ2ネガティブプロンプトは、直接的抑制によるものか、潜在空間における中和によるものか、オブジェクト削除をどのように達成しているか?
  • RQ3なぜネガティブプロンプトの早期適用は、逆誘発(逆活性化)と呼ばれる逆説的対象生成を引き起こすことがあるのか?
  • RQ4拡散プロセスの初期段階でネガティブプロンプトを適用すると、どのような構造的リスクが生じるか?
  • RQ5単純で侵襲的でないタイミング戦略が、モデル変更なしにインpainting性能を顕著に向上させられるか?

主な発見

  • ネガティブプロンプトは顕著な遅延を示し、正のプロンプトが対応するコンテンツをレンダリングした後、初めて影響を及ぼす。
  • 削除は、潜在空間内で正のプロンプトノイズから負のプロンプトノイズを差し引く中和メカニズムによって実現される。
  • ネガティブプロンプトを早期に適用すると、「逆誘発(reverse activation)」が生じ、ノイズダイナミクスにおけるモーメンタムと誘導効果により、禁止されたオブジェクトが逆に生成されることがある。
  • ネガティブプロンプトを適用する最適タイミングは、逆拡散プロセスの中点であり、構造的歪みを最小限に抑える。
  • 提案手法は、Placesで最大87.20%の相対的除去成功率(RRSR)、Vatexで86.44%を達成し、GPT-4V類似度スコアは平均82.84%を記録した。
  • 人間評価により、画像類似度が高く、平均して比較率(CR)が88.45%に達し、強い知覚的忠実度を示した。
Figure 2 : Illustration: Reverse activation. Each column shows an image generated by applying negative prompts in some specific steps which is shown at the top of the picture. In these two examples, the diffusion process without applying a negative prompt does not produce the object mentioned in the
Figure 2 : Illustration: Reverse activation. Each column shows an image generated by applying negative prompts in some specific steps which is shown at the top of the picture. In these two examples, the diffusion process without applying a negative prompt does not produce the object mentioned in the

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。