Skip to main content
QUICK REVIEW

[論文レビュー] Describing Human Aesthetic Perception by Deeply-learned Attributes from Flickr.

Luming Zhang|arXiv (Cornell University)|May 25, 2016
Visual Attention and Saliency Detection参考文献 18被引用数 8
ひとこと要約

本論文では、Flickrの画像タグから局所的で解釈可能な視覚的美的特徴を学習する弱教師付き深層学習フレームワークを提案する。画像レベルのテキスト的特徴を畳み込みニューラルネットワーク(CNN)を介して画素レベルのパッチに投影することで、領域特有の魅力を捉え、美的順序付け、リtrieval、リターゲティングにおいて優れた性能を達成する。

ABSTRACT

Many aesthetic models in computer vision suffer from two shortcomings: 1) the low descriptiveness and interpretability of those hand-crafted aesthetic criteria (i.e., nonindicative of region-level aesthetics), and 2) the difficulty of engineering aesthetic features adaptively and automatically toward different image sets. To remedy these problems, we develop a deep architecture to learn aesthetically-relevant visual attributes from Flickr1, which are localized by multiple textual attributes in a weakly-supervised setting. More specifically, using a bag-ofwords (BoW) representation of the frequent Flickr image tags, a sparsity-constrained subspace algorithm discovers a compact set of textual attributes (e.g., landscape and sunset) for each image. Then, a weakly-supervised learning algorithm projects the textual attributes at image-level to the highly-responsive image patches at pixel-level. These patches indicate where humans look at appealing regions with respect to each textual attribute, which are employed to learn the visual attributes. Psychological and anatomical studies have shown that humans perceive visual concepts hierarchically. Hence, we normalize these patches and feed them into a five-layer convolutional neural network (CNN) to mimick the hierarchy of human perceiving the visual attributes. We apply the learned deep features on image retargeting, aesthetics ranking, and retrieval. Both subjective and objective experimental results thoroughly demonstrate the competitiveness of our approach.

研究の動機と目的

  • 従来の手作業で作成された美的基準における解釈可能性の欠如と領域特異性の欠如に対処すること。
  • 多様な画像セットにおいて自動的かつ適応的に美的特徴を学習すること。
  • 視覚的特徴の階層的処理を模倣することで、人間の美的認識をモデル化すること。
  • 美的予測、画像リターゲティング、リtrievalタスクにおける性能を向上させること。

提案手法

  • 頻度の高いFlickr画像タグのbag-of-words(BoW)表現を用いて、各画像ごとのテキスト的特徴を抽出する。
  • スパarsity制約付き部分空間アルゴリズムを適用し、コンactかつ意味のあるテキスト的特徴の集合(例:'landscape'、'sunset')を発見する。
  • 弱教師付き学習法を用いて、画像レベルのテキスト的特徴を高応答性を持つ画素レベルの画像パッチに投影する。
  • 局所的パッチを正規化し、それらを5層の畳み込みニューラルネットワーク(CNN)に供給することで、階層的視覚的認識をモデル化する。
  • 学習された深層特徴を、美的順序付け、リtrieval、画像リターゲティングなどの下流タスクに活用する。
  • 心理的および解剖学的証拠(視覚的処理の階層的性質)を基に、CNNアーキテクチャの設計をガイドする。

実験結果

リサーチクエスチョン

  • RQ1画像レベルのタグのみを用いて、弱教師付き学習が画素レベルで美的特徴を局所化するのに効果的であるか。
  • RQ2テキスト的特徴に基づく学習された深層特徴が、美的予測および画像理解タスクの性能をどの程度向上させるか。
  • RQ3局所化された視覚的特徴が、人間の魅力的な画像領域への注目とどの程度一致するか。
  • RQ4階層的CNNアーキテクチャが、人間の美的特徴認識を効果的に模倣できるか。
  • RQ5提案手法は、手作業による特徴工学なしに多様な画像セットに一般化できるか。

主な発見

  • 提案手法は、手作業で作成された特徴に依存する従来のモデルを上回る性能を示し、画像の美的順序付けにおいて競争力のある結果を達成した。
  • 学習された視覚的特徴は局所的かつ解釈可能であり、美的魅力に寄与する特定の画像領域を示している。
  • 弱教師付きパッチ局所化は、テキスト的特徴が示す人間の魅力的な領域への注目を的確に反映している。
  • 5層のCNNは、人間の視覚的認識の階層的性質を効果的にモデル化し、特徴表現を改善した。
  • 画像リターゲティングやリtrievalなどの下流タスクへの一般化が良好であり、堅牢性と適応性を示した。
  • 主観的および客観的評価により、ベースラインの美的モデルよりも提案手法の優位性が確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。