Skip to main content
QUICK REVIEW

[논문 리뷰] MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Kunchang Li, Yali Wang|arXiv (Cornell University)|2023. 11. 28.
Multimodal Machine Learning Applications인용 수 4
한 줄 요약

MVBench는 다양한 복잡한 작업을 통해 비디오 추론 모델을 평가하기 위해 설계된 종합적인 다중모달 비디오 이해 벤치마크를 소개한다. 이 벤치마크는 비디오 질의 응답, 동작 예측, 시간적 기준 설정 등의 과제를 평가하며, VideoChat2가 96.4%의 히트 비율과 50.1의 평균 점수를 기록해 벤치마크에서 최고의 성능을 보이며 최신 기술 수준의 성과를 입증한다.

ABSTRACT

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess spatial understanding in the static image tasks, while overlooking temporal understanding in the dynamic video tasks. To alleviate this issue, we introduce a comprehensive Multi-modal Video understanding Benchmark, namely MVBench, which covers 20 challenging video tasks that cannot be effectively solved with a single frame. Specifically, we first introduce a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, we enable the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition. Then, guided by the task definition, we automatically convert public video annotations into multiple-choice QA to evaluate each task. On one hand, such a distinct paradigm allows us to build MVBench efficiently, without much manual intervention. On the other hand, it guarantees evaluation fairness with ground-truth video annotations, avoiding the biased scoring of LLMs. Moreover, we further develop a robust video MLLM baseline, i.e., VideoChat2, by progressive multi-modal training with diverse instruction-tuning data. The extensive results on our MVBench reveal that, the existing MLLMs are far from satisfactory in temporal understanding, while our VideoChat2 largely surpasses these leading models by over 15% on MVBench. All models and data are available at https://github.com/OpenGVLab/Ask-Anything.

연구 동기 및 목표

  • 다양한 비디오 추론 과제를 통해 다중모달 비디오 이해 모델을 평가하기 위한 종합적인 벤치마크를 구축하기 위해.
  • 비디오 추론 시스템을 위한 표준화되고 다양한 과제들로 구성된 도전적인 평가 프로토콜의 부족을 보완하기 위해.
  • 통합된 다면적 평가 프레임워크를 통해 최신 기술 수준의 모델들을 체계적으로 비교할 수 있도록 하기 위해.
  • 비디오 내에서 복잡한 시각적, 언어적, 시간적 관계를 이해할 수 있는 능력을 갖춘 모델의 개발과 평가를 지원하기 위해.

제안 방법

  • 벤치마크는 비디오 질의 응답, 동작 예측, 시간적 기준 설정 등을 포함한 다양한 비디오 이해 과제로 구성되어 있다.
  • 모델 성능를 비교하기 위해 히트 비율 및 평균 점수와 같은 지표를 사용하는 표준화된 평가 프로토콜을 사용한다.
  • 모델의 일반화 능력을 시험하기 위해 복잡한 다단계 추론 요구 사항을 수반한 다양한 비디오 데이터를 포함한다.
  • 모델가 주어진 선택지 중에서 정답을 올바르게 선택할 수 있는 능력, 특히 추론 및 시간적 이해 능력을 중심으로 평가한다.
  • 향후 새로운 비디오 추론 모델의 평가를 지원할 수 있도록 확장 가능하고 스케일이 가능한 설계를 한다.

실험 결과

연구 질문

  • RQ1현재의 다중모달 비디오 모델들은 다양한 종합적인 비디오 추론 과제에서 얼마나 잘 성능을 내는가?
  • RQ2다양한 평가 지표에서 서로 다른 비디오 추론 모델들의 상대적 성능는 어떠한가?
  • RQ3통합된 벤치마크가 실제 세계의 비디오 이해 과제의 복잡성과 다양성을 효과적으로 포괄할 수 있는가?
  • RQ4다단계의 다중모달 추론을 요구하는 비디오 데이터에 직면했을 때 기존 모델들의 한계는 무엇인가?

주요 결과

  • VideoChat2는 96.4%의 최고 히트 비율과 50.1의 평균 점수를 기록하여 다른 모델들보다 뛰어난 성능을 보였다.
  • VideoChat는 히트 비율 78.2%와 평균 점수 22.8을 기록하여 평가 과제에서 중간 수준의 성능를 보였다.
  • VideoChatGPT는 히트 비율 64.6%와 평균 점수 22.0을 기록하여 VideoChat2에 비해 낮은 성능를 보였다.
  • 벤치마크는 최신 기술 수준의 모델들과 인간 수준의 이해 능력 사이에 뚜렷한 성능 격차가 있음을 드러내며, 향후 개선 여지가 많음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.