Skip to main content
QUICK REVIEW

[論文レビュー] Understanding Aesthetics with Language: A Photo Critique Dataset for Aesthetic Assessment

Daniel Vera Nieto, L. Celona|arXiv (Cornell University)|Jun 17, 2022
Visual Attention and Saliency Detection被引用数 7
ひとこと要約

本論文は、74,000枚の高解像度画像と220,000件のユーザー生成写真批評をペアで含む大規模なデータセットであるReddit Photo Critique Dataset (RPCD) を紹介する。本研究では、批評の感情極性を美的判断の代理指標として用いることを提案し、人間による美的評価と強い相関関係があることを示し、批評の感情極性で微調整されたビジョントランスフォーマーモデルが、先行手法を上回ることを示している。

ABSTRACT

Computational inference of aesthetics is an ill-defined task due to its subjective nature. Many datasets have been proposed to tackle the problem by providing pairs of images and aesthetic scores based on human ratings. However, humans are better at expressing their opinion, taste, and emotions by means of language rather than summarizing them in a single number. In fact, photo critiques provide much richer information as they reveal how and why users rate the aesthetics of visual stimuli. In this regard, we propose the Reddit Photo Critique Dataset (RPCD), which contains tuples of image and photo critiques. RPCD consists of 74K images and 220K comments and is collected from a Reddit community used by hobbyists and professional photographers to improve their photography skills by leveraging constructive community feedback. The proposed dataset differs from previous aesthetics datasets mainly in three aspects, namely (i) the large scale of the dataset and the extension of the comments criticizing different aspects of the image, (ii) it contains mostly UltraHD images, and (iii) it can easily be extended to new data as it is collected through an automatic pipeline. To the best of our knowledge, in this work, we propose the first attempt to estimate the aesthetic quality of visual stimuli from the critiques. To this end, we exploit the polarity of the sentiment of criticism as an indicator of aesthetic judgment. We demonstrate how sentiment polarity correlates positively with the aesthetic judgment available for two aesthetic assessment benchmarks. Finally, we experiment with several models by using the sentiment scores as a target for ranking images. Dataset and baselines are available (https://github.com/mediatechnologycenter/aestheval).

研究の動機と目的

  • 単一の美的スコアの限界を克服し、自然言語による批評を通じてより豊かで解釈可能なフィードバックを捉えること。
  • 実世界の写真共有コミュニティから得た、大規模かつ高解像度の画像-コメントペアデータセットを、美的評価研究のためのものとして構築すること。
  • 明示的な人間の評価が不要な状況でも、写真批評の感情極性が美的判断の信頼できる代理指標として機能するかを検証すること。
  • テキスト的批評を活用するマルチモーダルモデルの、画像の美的順序付けおよびキャプション生成における有効性を評価すること。
  • 今後の計算美的評価およびマルチモーダル推論分野の研究を支援するため、公開可能で拡張可能なデータセットとベースラインモデルを提供すること。

提案手法

  • 写真フィードバックに特化したRedditコミュニティから、74,000枚の画像と220,000件の写真批評を収集した。
  • コメントが存在しない、または画像が利用できない投稿を除外し、1段階目のコメントのみを保持するなど、前処理を実施した。
  • 自然言語処理技術を用いて批評から感情極性を抽出し、美的判断の代理指標とした。
  • ビジョントランスフォーマー(ViT)モデルを、批評の感情極性スコアを順序付けターゲットとして用いて訓練・評価し、画像の美的評価を実施した。
  • 美的評価と美的キャプション生成の両タスクを評価可能な統合評価フレームワークを設計した。
  • GitHubおよびZenodoを通じて、データセット、コード、ベースラインをクリエイティブ・コモンズ 4.0 ライセンスで公開し、今後の研究や拡張を可能にした。

実験結果

リサーチクエスチョン

  • RQ1明示的な人間の評価がなくても、ユーザーが生成した写真批評の感情極性は、画像の美的質を信頼性高く予測できるか?
  • RQ2従来の美的スコアではなく、批評の感情極性で微調整されたビジョンモデルの性能は、従来手法と比べてどのように異なるか?
  • RQ3単一の美的スコアに比べて、美的批評コメントは、計算的美的評価においてより情報量の多い信号を提供するか?
  • RQ4感情極性を理解するように学習したモデルは、美的画像キャプション生成などの他のタスクにも一般化可能か?
  • RQ5RPCDの規模と解像度は、既存のデータセットと比較して、マルチモーダル学習における実用性においてどの程度優れているか?

主な発見

  • 2つのベンチマークデータセットにおいて、写真批評の感情極性は人間による美的スコアと強い正の相関関係を示した。
  • 批評の感情極性で微調整されたビジョントランスフォーマーは、画像の美的評価において最先端の性能を達成した。
  • 批評の感情極性から美的に注意を向ける特徴を学習させることで、単に意味的特徴のみを用いるモデルに比べて顕著な性能向上が得られた。
  • 同じアーキテクチャを用いても、美的順序付けとキャプション生成といった異なる目的で訓練したモデルは、異なる性能特性を示し、タスク固有の最適化の重要性を示した。
  • RPCDは、美的評価を目的とした画像-コメントペアのうち、最大級かつ最高解像度のコレクションであり、既存のデータセットと比較してコメントがより長く、より情報量に富んでいる。
  • データセットとベースラインは完全に再現可能であり、公開されており、今後の研究や自動パイプラインによる拡張が可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。