Skip to main content
QUICK REVIEW

[論文レビュー] EmoSet: A Large-scale Visual Emotion Dataset with Rich Attributes

Jingyuan Yang, Qirui Huang|arXiv (Cornell University)|Jul 16, 2023
Advanced Computing and Algorithms被引用数 5
ひとこと要約

EmoSetは、明るさ、彩度、シーンタイプ、オブジェクトクラス、顔の表情、人間の行動といった、心理学的に根拠を持つ豊富な属性を備えた、118,102枚の画像にアノテーションが施された、初めての大規模な視覚的感情データセットを提供する。属性に配慮したモデリングにより、視覚的感情認識で優れた性能を達成しており、EmoSetの事前学習から微調整した場合、FIデータセットで20.09%の精度低下を示し、優れた汎化性能と解釈可能性の利点を示している。

ABSTRACT

Visual Emotion Analysis (VEA) aims at predicting people's emotional responses to visual stimuli. This is a promising, yet challenging, task in affective computing, which has drawn increasing attention in recent years. Most of the existing work in this area focuses on feature design, while little attention has been paid to dataset construction. In this work, we introduce EmoSet, the first large-scale visual emotion dataset annotated with rich attributes, which is superior to existing datasets in four aspects: scale, annotation richness, diversity, and data balance. EmoSet comprises 3.3 million images in total, with 118,102 of these images carefully labeled by human annotators, making it five times larger than the largest existing dataset. EmoSet includes images from social networks, as well as artistic images, and it is well balanced between different emotion categories. Motivated by psychological studies, in addition to emotion category, each image is also annotated with a set of describable emotion attributes: brightness, colorfulness, scene type, object class, facial expression, and human action, which can help understand visual emotions in a precise and interpretable way. The relevance of these emotion attributes is validated by analyzing the correlations between them and visual emotion, as well as by designing an attribute module to help visual emotion recognition. We believe EmoSet will bring some key insights and encourage further research in visual emotion analysis and understanding. Project page: https://vcc.tech/EmoSet.

研究の動機と目的

  • 豊富なアノテーションを備えた大規模で多様かつバランスの取れた視覚的感情データセットの不足に対処すること。
  • 心理学的理論に基づく記述可能な視覚的属性を導入することで、視覚的感情分析における感情ギャップを埋めること。
  • 補助的属性の監視を通じて、より解釈可能で頑健な視覚的感情認識を可能にすること。
  • 感情認識を超えて、感情的刺激のより深い理解を促進すること。
  • 広範な画像-テキストおよび属性アノテーションのおかげで、マルチモーダルおよび弱教師付き学習を支援すること。

提案手法

  • Mikelsの感情モデルを用いて810の感情キーワードを用いて取得した330万枚の画像から、EmoSet-3.3Mを構築した。
  • 8つの感情カテゴリと6つの視覚的属性を持つ118,102枚の画像を収集し、人間によるアノテーションを実施した。
  • 明るさ、シーンタイプ、顔の表情からの感情関連特徴を抽出・統合するためのマルチブランチ属性モジュールを設計した。
  • t-SNE可視化を用いて、属性特徴が判別性があり、感情に配慮した表現を学習していることを確認した。
  • ImageNet事前学習の有無にかかわらず、ResNet-50を用いてFIおよびArtphotoデータセットでクロスデータセット一般化を評価した。
  • 相関分析とアブレーションスタディを通じて、属性の関連性を感情認識性能の観点から検証した。

実験結果

リサーチクエスチョン

  • RQ1明るさや顔の表情といった豊富な視覚的属性は、視覚的刺激における特定の感情カテゴリとどのように相関しているか?
  • RQ2標準的な認識手法と比較して、属性に配慮したモデリングは視覚的感情認識性能をどの程度向上させるか?
  • RQ3EmoSetで事前学習したモデルは、FI や Artphoto といった他の視覚的感情データセットにどの程度一般化するか?
  • RQ4属性特徴は、異なる感情状態の間で意味的に分離された形で可視化できるか?
  • RQ5スケール、多様性、バランス、アノテーションの豊富さの観点から、EmoSetは既存のデータセットと比べてどの程度優れているか?

主な発見

  • EmoSet-118Kには118,102枚の人間によるアノテーションが施された画像が含まれており、最大の既存データセット(FI)の5倍以上大きく、8つの感情カテゴリにわたるクラス分布がバランスが取れている。
  • 属性モジュールにより、判別性があり感情に関連する特徴を学習することで、視覚的感情認識が向上した。t-SNE可視化では、感情ごとに明確なクラスタリングが観察された。
  • EmoSetで事前学習したモデルは、FIでテストした際の精度低下が20.09%にとどまり、FIから微調整した場合の26.98%の低下よりも小さく、EmoSetからの汎化性能が優れていることを示している。
  • 属性特徴としての「廃墟(ruin)」と「遊び部屋(playroom)」は、t-SNEプロット上で空間的に分離されており、それぞれ悲しみと楽しさに対応しており、属性の関連性が裏付けられた。
  • EmoSetは、ソーシャルコンテンツや芸術的コンテンツを含む多様な画像ソースを有しているため、芸術的画像(Artphoto)に対しても良好に一般化し、EmoSetで学習したモデルがFIで学習したモデルを上回る性能を示した。
  • 画像-テキストおよび属性アノテーションの豊富さのおかげで、弱教師付き学習、視覚言語モデリング、今後の視覚的感情生成や編集に関する研究を支援できる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。