[論文レビュー] MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild
本稿では、10,045本の動画・音声クリップに複合感情および感情的行動のラベルが付与された大規模かつ現実世界を想定したマルチモーダルデータベースMAFWを紹介する。また、複数モodal間での表情変化関係を活用する新規のトランスフォーマー基盤手法T-ESFLを提案し、動的顔面表情認識の性能を向上させ、単一およびマルチモーダルFERタスクの最先端手法を上回る結果を得た。
Dynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain only one modality, e.g., videos. The monotonous labels and modality cannot accurately imitate human emotions and fulfill applications in the real world. In this paper, we propose MAFW, a large-scale multi-modal compound affective database with 10,045 video-audio clips in the wild. Each clip is annotated with a compound emotional category and a couple of sentences that describe the subjects' affective behaviors in the clip. For the compound emotion annotation, each clip is categorized into one or more of the 11 widely-used emotions, i.e., anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment. To ensure high quality of the labels, we filter out the unreliable annotations by an Expectation Maximization (EM) algorithm, and then obtain 11 single-label emotion categories and 32 multi-label emotion categories. To the best of our knowledge, MAFW is the first in-the-wild multi-modal database annotated with compound emotion annotations and emotion-related captions. Additionally, we also propose a novel Transformer-based expression snippet feature learning method to recognize the compound emotions leveraging the expression-change relations among different emotions and modalities. Extensive experiments on MAFW database show the advantages of the proposed method over other state-of-the-art methods for both uni- and multi-modal FER. Our MAFW database is publicly available from https://mafw-database.github.io/MAFW.
研究の動機と目的
- 既存の動的顔面表情データベースが基本的かつ互いに排他的な感情のみを用いているという制限に対処する。
- 現実世界の感情の複雑さを反映した、複合感情ラベルが付与された大規模で現実世界を想定したマルチモーダルデータベースを構築する。
- 視覚的、音声的、文書的モダリティ間の表情変化ダイナミクスをモデル化することで、複合感情認識のための堅牢な手法を開発する。
- 高品質で文脈豊富なマルチモーダルデータを提供することで、感情コンピューティング分野における新たな研究を可能にする。
提案手法
- 10,045本の動画・音声クリップと20,000件のテキスト的感情的キャプションを備えた大規模で現実世界を想定したマルチモーダルデータベースMAFWを提案する。
- 信頼性の低いラベルをフィルタリングするために期待最大化(EM)アルゴリズムを用い、11種類の単一ラベルおよび32種類のマルチラベル感情カテゴリを特定した。
- 複数モダリティ間での表情変化関係をモデル化する新規のトランスフォーマー基盤の表現スニペット特徴学習(T-ESFL)手法を導入する。
- 視覚的特徴学習にはスニペットベースのトランスフォーマーと空間周波数順序再構成(SSOR)を、音声にはResNet_LSTMを、テキストにはDPCNNを採用する。
- 各モダリティ固有の特徴を連結することでマルチモーダル感情表現を構築する。
- 交差エントロピー損失とスニペット順序再構成損失を組み合わせた共同目的関数を最適化することで、時間的ダイナミクスのモデリングを強化する。
実験結果
リサーチクエスチョン
- RQ1複合感情ラベルが付与された大規模で現実世界を想定したマルチモーダルデータベースは、動的顔面表情認識の現実性と複雑さを向上させることができるか?
- RQ2提案手法T-ESFLは、単一およびマルチモーダル設定の両方において、最先端のアプローチと比較して、複合感情認識においてどの程度有効であるか?
- RQ3複数モダリティ間での表情変化関係は、動的FERにおける認識性能をどの程度向上させるか?
- RQ4MAFWデータベースは、高品質で文脈豊富な記述を提供するため、動画感情キャプションなどの後続タスクを支援できるか?
主な発見
- 提案手法T-ESFLは、MAFWデータセット上で単一およびマルチモーダル動的顔面表情認識の両方で最先端の性能を達成し、既存手法を上回った。
- 汎用モデルを用いた動画感情キャプション生成では、BLEU-4が9.09、METEORが15.49、CIDErが23.40を達成し、MAFWが自然言語生成タスクに有用であることを示した。
- EMベースのフィルタリングプロセスにより、ラベルノイズが著しく低減され、信頼性の高い11種類の単一ラベルおよび32種類のマルチラベル感情カテゴリが得られた。
- 定性的な分析から、モデルが生成した感情的キャプションが正解ラベルとよく一致しており、微細な感情的行動および文脈的手がかりを的確に捉えていることが確認された。
- MAFWデータベースはhttps://mafw-database.github.io/MAFWで公開されており、感情コンピューティング分野における広範な研究を促進する。
- データセットは最小限の人口統計的バイアスを示しており、性別の統計はCelebAから推定され、データ分布分析の目的でのみ使用され、モデル学習には使用されていない。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。