[論文レビュー] VideoStory Embeddings Recognize Events when Examples are Scarce
この論文では、記述性とマルチモodal予測可能性の両方をバランスさせる共同目的関数を用いて、動画特徴量とテキスト記述を統合された意味的空間に埋め込むVideoStoryという動画表現学習手法を提案する。ウェブスケールの動画-テキストペアを活用し、相関する用語を最適化することで、TRECVIDおよびColumbia Consumer Videosデータセットにおいて、少サンプルおよびゼロサンプルのイベント認識の両面で最先端の性能を達成した。
This paper aims for event recognition when video examples are scarce or even completely absent. The key in such a challenging setting is a semantic video representation. Rather than building the representation from individual attribute detectors and their annotations, we propose to learn the entire representation from freely available web videos and their descriptions using an embedding between video features and term vectors. In our proposed embedding, which we call VideoStory, the correlations between the terms are utilized to learn a more effective representation by optimizing a joint objective balancing descriptiveness and predictability.We show how learning the VideoStory using a multimodal predictability loss, including appearance, motion and audio features, results in a better predictable representation. We also propose a variant of VideoStory to recognize an event in video from just the important terms in a text query by introducing a term sensitive descriptiveness loss. Our experiments on three challenging collections of web videos from the NIST TRECVID Multimedia Event Detection and Columbia Consumer Videos datasets demonstrate: i) the advantages of VideoStory over representations using attributes or alternative embeddings, ii) the benefit of fusing video modalities by an embedding over common strategies, iii) the complementarity of term sensitive descriptiveness and multimodal predictability for event recognition without examples. By it abilities to improve predictability upon any underlying video feature while at the same time maximizing semantic descriptiveness, VideoStory leads to state-of-the-art accuracy for both few- and zero-example recognition of events in video.
研究の動機と目的
- ラベル付き例が不足または存在しない動画におけるイベント認識を解決すること。
- 動画特徴量から予測可能で、かつ意味的に記述的である動画表現を開発すること。
- 自由に入手可能なウェブ動画とそのテキスト記述から学習することで、少サンプルおよびゼロサンプル設定における一般化性能を向上させること。
- 外見、動き、音声といった複数のモodalを統合された意味的埋め込みに統合し、より高い予測性能を実現すること。
- 動画例が一切不要なテキストクエリからの正確なイベント認識を可能にすること。
提案手法
- 外見、動き、音声特徴量のマルチモーダル予測可能性損失を用いて、動画特徴量と語彙ベクトルの間で共同埋め込みを学習すること。
- 語彙が動画をどれだけ適切に記述できるか(記述性)と、動画特徴量が語彙をどれだけ適切に予測できるか(予測可能性)の両方をバランスさせる共同目的関数を最適化すること。
- 外見、動き、音声特徴量を共有意味的空間に統合するVideoStoryFという変種を導入し、予測性能を向上させること。
- ゼロサンプル認識における精度を向上させるために、語彙に敏感な記述性損失を用いるVideoStory0という変種を提案すること。
- 手動でアノテートされた属性に依存せずに、ウェブから収集した動画-テキストペアを用いて埋め込みを事前学習すること。
- 任意の下位の動画特徴量(例:CNN、改善版密なトレイジェクトリ)にVideoStory埋め込みを適用し、その予測性能を向上させること。
実験結果
リサーチクエスチョン
- RQ1手動でアノテートされた属性に依存せずに、ウェブ動画とその記述から意味的動画表現を効果的に学習できるか?
- RQ2記述性とマルチモーダル予測可能性の共同最適化が、少サンプルおよびゼロサンプル設定におけるイベント認識を向上させるか?
- RQ3共有埋め込みを用いて外見、動き、音声特徴量を統合することで、従来の特徴量統合手法を上回る性能が得られるか?
- RQ4語彙に敏感な記述性損失は、テキストクエリからのゼロサンプルイベント認識の精度向上にどの程度効果的か?
- RQ5VideoStoryは、少サンプルおよびゼロサンプルのイベント検出において、複数のベンチマークで最先端の性能を達成できるか?
主な発見
- TRECVID 2013の少サンプルテストセットにおいて、VideoStoryは平均平均精度(mAP)37.1%を達成し、新たな最先端の成績を記録した。
- ゼロサンプル設定では、VideoStoryはmAP 20.0%を達成し、以前の最高成績である18.3%を上回った。
- 外見、動き、音声特徴量を統合するVideoStoryFは、予測可能性を向上させ、従来の統合手法よりも高い精度を実現した。
- 語彙に敏感な記述性損失を用いるVideoStory0は、テキストクエリのみから正確な認識を可能にし、強力なゼロショット一般化性能を示した。
- 語彙に敏感な記述性損失とマルチモーダル予測可能性損失の組み合わせは相補的であり、ゼロサンプル認識の性能を顕著に向上させた。
- VideoStoryは、すべての下位の動画特徴量(例:CNN、I3D、音声)において予測可能性を向上させつつ、高い意味的記述性を維持し、最先端の結果を達成した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。