Skip to main content
QUICK REVIEW

[論文レビュー] MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Kunchang Li, Yali Wang|arXiv (Cornell University)|Nov 28, 2023
Multimodal Machine Learning Applications被引用数 4
ひとこと要約

MVBenchは、多様で複雑なタスクをカバーするように設計された包括的なマルチモーダル動画理解ベンチマークを導入した。このベンチマークは、動画質問応答、行動予測、時系列グランドトゥングの3つの分野でモデルを評価し、VideoChat2が96.4%のヒットレートと50.1の平均スコアを記録し、ベンチマーク上で最先端の結果を示した。

ABSTRACT

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess spatial understanding in the static image tasks, while overlooking temporal understanding in the dynamic video tasks. To alleviate this issue, we introduce a comprehensive Multi-modal Video understanding Benchmark, namely MVBench, which covers 20 challenging video tasks that cannot be effectively solved with a single frame. Specifically, we first introduce a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, we enable the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition. Then, guided by the task definition, we automatically convert public video annotations into multiple-choice QA to evaluate each task. On one hand, such a distinct paradigm allows us to build MVBench efficiently, without much manual intervention. On the other hand, it guarantees evaluation fairness with ground-truth video annotations, avoiding the biased scoring of LLMs. Moreover, we further develop a robust video MLLM baseline, i.e., VideoChat2, by progressive multi-modal training with diverse instruction-tuning data. The extensive results on our MVBench reveal that, the existing MLLMs are far from satisfactory in temporal understanding, while our VideoChat2 largely surpasses these leading models by over 15% on MVBench. All models and data are available at https://github.com/OpenGVLab/Ask-Anything.

研究の動機と目的

  • 多様な動画推論タスクにわたるマルチモーダル動画理解モデルの包括的評価のためのベンチマークを確立すること。
  • 動画推論システムのための標準的で多様かつ挑戦的な評価プロトコルの不足を是正すること。
  • 統一的で多面的な評価フレームワーク上で、最先端モデルの体系的比較を可能にすること。
  • 動画内の複雑な視覚的・言語的・時系列的関係を理解できるモデルの開発と評価を支援すること。

提案手法

  • ベンチマークは、動画質問応答、行動予測、時系列グランドトゥングを含む多様な動画理解タスクで構成される。
  • ヒットレートや平均スコアといったメトリクスを用いた標準化された評価プロトコルを採用し、モデルのパフォーマンスを比較する。
  • モデルの汎化能力をテストするため、複雑で複数ターンの推論を要する多様な動画データを含む評価が実施される。
  • モデルが選択肢の中から正しい答えを特定する能力、特に推論と時系列理解に焦点を当てる。
  • スケーラブルで拡張可能な設計となっており、将来の動画推論モデルの評価をサポートする。

実験結果

リサーチクエスチョン

  • RQ1現在のマルチモーダル動画モデルは、多様で包括的な動画推論タスクのセットでどの程度のパフォーマンスを示すか?
  • RQ2複数の評価メトリクスにわたって、異なる動画推論モデルの相対的なパフォーマンスはいかなるものか?
  • RQ3統一されたベンチマークは、現実世界の動画理解課題の複雑さと多様性を効果的に捉えることができるか?
  • RQ4複数ターンのマルチモーダル推論を伴う動画データに対して、既存のモデルにどのような限界があるか?

主な発見

  • VideoChat2は96.4%の最高ヒットレートと50.1の平均スコアを記録し、ベンチマーク全体で他のモデルを上回った。
  • VideoChatはヒットレート78.2%、平均スコア22.8を達成し、評価タスクで中程度のパフォーマンスを示した。
  • VideoChatGPTはヒットレート64.6%、平均スコア22.0を記録し、VideoChat2に比べて低いパフォーマンスを示した。
  • ベンチマークは、最先端モデルと人間レベルの理解の間で顕著なパフォーマンス格差を明らかにし、改善の余地が広く存在することを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。