Skip to main content
QUICK REVIEW

[論文レビュー] Understanding Chinese Video and Language via Contrastive Multimodal Pre-Training

Chenyi Lei, Shixian Luo|arXiv (Cornell University)|Apr 19, 2021
Multimodal Machine Learning Applications被引用数 4
ひとこと要約

本稿では、複数モodalの事前学習フレームワークVictorを提案する。これは中国語の動画・言語理解を目的とした対照的マルチモーダル事前学習フレームワークであり、複雑な時空間的および順序的関係を捉えるために、新規の再構築的および対照的プロキシタスクを導入している。大規模なAlivOL-10Mデータセットを用いて訓練されたVictorは、動画検索、分類、推薦、キャプション生成といった複数の下流タスクで最先端の性能を達成している。

ABSTRACT

The pre-trained neural models have recently achieved impressive performances in understanding multimodal content. However, it is still very challenging to pre-train neural models for video and language understanding, especially for Chinese video-language data, due to the following reasons. Firstly, existing video-language pre-training algorithms mainly focus on the co-occurrence of words and video frames, but ignore other valuable semantic and structure information of video-language content, e.g., sequential order and spatiotemporal relationships. Secondly, there exist conflicts between video sentence alignment and other proxy tasks. Thirdly, there is a lack of large-scale and high-quality Chinese video-language datasets (e.g., including 10 million unique videos), which are the fundamental success conditions for pre-training techniques. In this work, we propose a novel video-language understanding framework named VICTOR, which stands for VIdeo-language understanding via Contrastive mulTimOdal pRe-training. Besides general proxy tasks such as masked language modeling, VICTOR constructs several novel proxy tasks under the contrastive learning paradigm, making the model be more robust and able to capture more complex multimodal semantic and structural relationships from different perspectives. VICTOR is trained on a large-scale Chinese video-language dataset, including over 10 million complete videos with corresponding high-quality textual descriptions. We apply the pre-trained VICTOR model to a series of downstream applications and demonstrate its superior performances, comparing against the state-of-the-art pre-training methods such as VideoBERT and UniVL. The codes and trained checkpoints will be publicly available to nourish further developments of the research community.

研究の動機と目的

  • 事前学習に適した大規模かつ高品質な中国語の動画・言語データセットの不足に対処する。
  • 動画とテキストにおける時空間的および順序的関係を十分に活用できない既存の動画・言語事前学習モデルの限界を克服する。
  • 対照的学習を導入することで、動画文書の対応付けと他のプロキシタスクとの矛盾を解消する。
  • 中国語の下流タスクにおけるマルチモーダル表現学習を強化する、堅牢な事前学習フレームワークを開発する。
  • 検索、分類、推薦、キャプション生成を含む多様な下流応用において、提案手法の有効性を実証する。

提案手法

  • エンコーダ・デコーダ型のTransformerベースのフレームワークであるVictorを提案する。
  • 順序構造を捉えるために、新規の再構築的プロキシタスク(マスクドフレーム順序モデリング:MFOM、マスクド文書順序モデリング:MSOM)を導入する。
  • 耐性を高め、不一致ペアの干渉を低減するため、二重の動画・テキスト対応(dual-VSA)を含む対照的プロキシタスクを設計する。
  • 高品質なテキスト記述を備えた1000万件を超える動画を含む大規模な中国語動画・言語データセット、AlivOL-10Mを活用する。
  • 再構築的および対照的目的を組み合わせて訓練することで、マルチモーダル表現学習を統合的に最適化する。
  • 分類タスクには[CLS]トークンの埋め込みを、キャプション生成タスクにはデコーダベースの生成を用いて、下流タスクでVictorを微調整する。
Figure 1. General proxy tasks that are widely used in existing video-language pre-training methods.
Figure 1. General proxy tasks that are widely used in existing video-language pre-training methods.

実験結果

リサーチクエスチョン

  • RQ1不一致の動画・テキストペアからの干渉を低減することで、対照的学習がマルチモーダル表現学習を向上させられるか?
  • RQ2MFOMやMSOMのような新規プロキシタスクは、動画と言語における順序的および構造的関係を捉える能力をどの程度向上させるか?
  • RQ3大規模かつ高品質な中国語動画・言語データセットで事前学習することで、既存手法と比較して下流タスクの性能がどの程度向上するか?
  • RQ4提案されたフレームワークは、VideoBERT や UniVL といった最先端モデルを、多様な動画・言語タスクで上回ることができるか?
  • RQ5時空間的関係(Intra-および Inter-MFMを介して)をモデル化することは、分類および検索タスクにどの程度寄与するか?

主な発見

  • 動画検索タスクにおいて、VictorはR@10が86.7%を達成し、UniVL や VideoBERT を含むすべてのベースラインを上回った。
  • マルチレベルの動画分類では、LeafCateタスクで92.1%のTop-1正答率、TopCateタスクで89.3%を達成し、M3 や M4 を上回った。
  • 動画推薦タスクでは、AUCが0.892、HR@5が0.721を達成し、DINベースラインおよび他の事前学習モデルを顕著に上回った。
  • マルチモーダル動画キャプション生成では、BLEU-4が38.2、Meteorが28.9、Ciderが124.1を達成し、すべての指標でUniVLを上回った。
  • アブレーションスタディの結果、MFOMおよびMSOMは順序的推論を顕著に向上させた一方、dual-VSAおよびMFMは検索および分類の正答率向上に寄与した。
  • AlivOL-10Mデータセット全体(1000万件)で学習させた場合、10万件のサブセットで学習させた場合よりも顕著に優れた結果が得られ、大規模データの重要性を裏付けた。
Figure 2. An example of the videos in Alivol -10M dataset. Besides the high-resolution video frames and human-created title, the video also contains a long-text abstract and a related e-commerce product with images. Each video has three types of categories: plot category, coarse-grained product cate
Figure 2. An example of the videos in Alivol -10M dataset. Besides the high-resolution video frames and human-created title, the video also contains a long-text abstract and a related e-commerce product with images. Each video has three types of categories: plot category, coarse-grained product cate

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。