Skip to main content
QUICK REVIEW

[論文レビュー] Style-Aware Normalized Loss for Improving Arbitrary Style Transfer

Jiaxin Cheng, Ayush Jaiswal|arXiv (Cornell University)|Apr 18, 2021
Generative Adversarial Networks and Image Synthesis参考文献 50被引用数 4
ひとこと要約

本稿では、任意のスタイル転送(AST)における不均衡なスタイル適合性(IST)を解消するため、スタイルに適応した正規化損失を提案する。トレーニング中にスタイル損失を等重みで扱うことで、モデルが画像を過剰にまたは不十分にスタイル化してしまう問題に対処する。理論的境界を導出し、スタイルに適応した正規化を導入することで、4つのASTモデル全体で騙し率に111%の相対的向上が得られ、人間の好みは98%向上した。

ABSTRACT

Neural Style Transfer (NST) has quickly evolved from single-style to infinite-style models, also known as Arbitrary Style Transfer (AST). Although appealing results have been widely reported in literature, our empirical studies on four well-known AST approaches (GoogleMagenta, AdaIN, LinearTransfer, and SANet) show that more than 50% of the time, AST stylized images are not acceptable to human users, typically due to under- or over-stylization. We systematically study the cause of this imbalanced style transferability (IST) and propose a simple yet effective solution to mitigate this issue. Our studies show that the IST issue is related to the conventional AST style loss, and reveal that the root cause is the equal weightage of training samples irrespective of the properties of their corresponding style images, which biases the model towards certain styles. Through investigation of the theoretical bounds of the AST style loss, we propose a new loss that largely overcomes IST. Theoretical analysis and experimental results validate the effectiveness of our loss, with over 80% relative improvement in style deception rate and 98% relatively higher preference in human evaluation.

研究の動機と目的

  • 任意のスタイル転送(AST)モデルにおける不均衡なスタイル適合性(IST)の根本的原因を解明すること。
  • 多様なスタイルにわたるサンプルごとの損失重みが等しく設定されていることで、モデルが特定のスタイルに偏ってしまうことの特定。
  • 理論的に裏付けられた、スタイルに適応した損失関数を提案し、スタイル固有の特性に基づいてトレーニング損失を正規化すること。
  • 4つの最先端のASTモデルを用いた広範なベンチマークと人間評価を通じて、新しい損失の妥当性を検証すること。
  • 損失最適化と人間のスタイル化品質への認識の整合性を高めること。

提案手法

  • 著者らは、従来のグラム行列に基づくスタイル損失の理論的境界を分析し、さまざまなスタイル特性下での期待値を導出する。
  • 各スタイル画像の固有の性質に基づいて、サンプルごとのスタイル損失を再重み付けする新しいスタイルに適応した正規化損失を提案する。
  • スタイル画像の特徴統計から得られる学習または推定されたスケール要因を用いて、各サンプルごとのスタイル損失を正規化する。
  • アーキテクチャの変更なしに既存のASTフレームワークに統合可能であり、即座に適用可能な改善が可能である。
  • 理論的分析により、新しい損失は人間の認識と正の相関があることが示され、従来の損失とは異なり、人間の評価スコアを反映していない。
  • ImageNetおよびPBN/DTDデータセットを用いた騙し率と人間評価の研究を通じて、アプローチを評価する。
Figure 2 : Distribution of classic Gram matrix-based style losses for four AST methods [ 14 , 19 , 29 , 37 ] . Smaller loss does not guarantee better style transfer (left two images) while high quality transferred images can have larger style losses (middle two images), with over-stylized images cou
Figure 2 : Distribution of classic Gram matrix-based style losses for four AST methods [ 14 , 19 , 29 , 37 ] . Smaller loss does not guarantee better style transfer (left two images) while high quality transferred images can have larger style losses (middle two images), with over-stylized images cou

実験結果

リサーチクエスチョン

  • RQ1最先端のASTモデルが多様なスタイルにわたって一般化できず、しばしば未適応または過剰にスタイル化された出力を生成する理由は何か?
  • RQ2従来のスタイル損失は、スタイル化品質の人間認識をどのように反映していないか?
  • RQ3トレーニング中に異なるスタイル画像間でスタイル適合性(IST)の不均衡が生じる原因は何か?
  • RQ4理論的に裏付けられた、スタイルに適応した損失の正規化により、トレーニングのバランスと出力品質が向上するか?
  • RQ5提案された損失は、複数のASTモデルおよび騙し率や人間の好みといったメトリクスにおいて有効か?

主な発見

  • 提案されたスタイルに適応した正規化損失は、全テスト対象のASTモデルで騙し率に111%の相対的向上を達成し、LinearTransferでは最大111%の向上を示した。
  • 人間評価では、新しい損失でスタイル化された画像に対して55.9%の好みが得られ、従来の損失と比較して98%の相対的増加を示した。
  • 新しい損失は人間の認識と正の相関があるが、従来の損失は人間の評価スコアを反映していない。
  • この手法はASTにおける未適応および過剰なスタイル化の問題を効果的に軽減し、モデル全体で不適切な出力を50%未満に削減した。
  • 理論的分析により、新しい損失はスタイル固有の特徴分布の分散を考慮することで、真のスタイル転送品質をよりよく反映していることが確認された。
  • この損失は普遍的に適用可能であり、GoogleMagenta、AdaIN、LinearTransfer、SANetの4つの異なるASTモデルで性能向上を達成した。
Figure 3 : Statistics of human perception of stylization quality as assessed in Study II.
Figure 3 : Statistics of human perception of stylization quality as assessed in Study II.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。