Skip to main content
QUICK REVIEW

[論文レビュー] ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration

Yuhao Cui, Yu Zhou|arXiv (Cornell University)|Aug 16, 2021
Multimodal Machine Learning Applications参考文献 52被引用数 4
ひとこと要約

ROSITAは、視覚・言語事前学習手法を提案し、統合されたシーングラフにクロスモodalおよびイントラモーダル知識を統合することで、微細な意味的整合性を向上させる。構造的知識マスキング(SKM)戦略を新たに導入し、マスクされた言語および領域モデリングを改善する。6つのベンチマークデータセットで3つの視覚・言語タスクにおいて最先端の性能を達成し、より強固で正確なクロスモーダル整合性を実現することで、既存手法を上回る。

ABSTRACT

Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning fine-grained semantic alignments between image-text pairs plays a key role in their approaches. Nevertheless, most existing VLP approaches have not fully utilized the intrinsic knowledge within the image-text pairs, which limits the effectiveness of the learned alignments and further restricts the performance of their models. To this end, we introduce a new VLP method called ROSITA, which integrates the cross- and intra-modal knowledge in a unified scene graph to enhance the semantic alignments. Specifically, we introduce a novel structural knowledge masking (SKM) strategy to use the scene graph structure as a priori to perform masked language (region) modeling, which enhances the semantic alignments by eliminating the interference information within and across modalities. Extensive ablation studies and comprehensive analysis verifies the effectiveness of ROSITA in semantic alignments. Pretrained with both in-domain and out-of-domain datasets, ROSITA significantly outperforms existing state-of-the-art VLP methods on three typical vision-and-language tasks over six benchmark datasets.

研究の動機と目的

  • 既存の視覚・言語事前学習(VLP)手法が画像・テキストペアからの内在的知識を十分に活用できないという制限に対処すること。
  • 画像領域とテキスト語の間の微細な意味的整合性を向上させること。これは、下流のV+Lタスクにおいて重要である。
  • 画像とテキスト間のクロスモーダル知識(クロスモーダル知識)と、各モーダル内でのイントラモーダル知識(イントラモーダル知識)を統合的なフレームワークで統合すること。
  • 構造的知識を活用するより効果的な事前学習を可能にする、堅牢なマスキング戦略を開発すること。
  • 多様な視覚・言語ベンチマークにおいて一貫した性能向上を示すこと。

提案手法

  • アンカーオブジェクトを中心にした統合されたシーングラフを構築し、クロスモーダル関係(例:'grass' ↔ 'steppe')とインフラモーダル関係(例:空間的に関連する領域、文脈的に関連する語)を符号化する。
  • アンカーオブジェクトと、シーングラフ内の構造的役割に基づいて選択的にマスクされる関連知識エントリをマスクする、画期的な構造的知識マスキング(SKM)戦略を導入する。
  • 各知識エントリごとに独立したマスキング確率を用いることで、情報マスキングに対する微細な制御を可能にし、一様確率戦略と比較して耐性を向上させる。
  • 既存のVLPモデルで用いられる標準的なマスク言語(領域)モデリング目的にSKMをスムーズに統合する。
  • 統合されたシーングラフとSKMを用いて、ドメイン内およびドメイン外のデータセットでモデルを事前学習し、クロスモーダル整合性を強化する。
Figure 1 . Schematic of the knowledge integration strategies of three VLP methods, i.e. , OSCAR (Li et al . , 2020b ) , ERNIE-ViL (Yu et al . , 2021 ) , and our ROSITA. OSCAR and ERNIE-ViL only exploit the intra-modal knowledge from the image and text modalities, respectively. In contrast, ROSITA si
Figure 1 . Schematic of the knowledge integration strategies of three VLP methods, i.e. , OSCAR (Li et al . , 2020b ) , ERNIE-ViL (Yu et al . , 2021 ) , and our ROSITA. OSCAR and ERNIE-ViL only exploit the intra-modal knowledge from the image and text modalities, respectively. In contrast, ROSITA si

実験結果

リサーチクエスチョン

  • RQ1統合されたシーングラフにクロスモーダルおよびインフラモーダル知識を統合することで、視覚・言語の意味的整合性が向上するか?
  • RQ2提案された構造的知識マスキング(SKM)戦略は、標準的なマスキング手法と比較して、整合性学習をどのように向上させるか?
  • RQ3各知識エントリごとに独立したマスキング確率を使用することで、同一の固定確率と比較して一般化性能が向上するか?
  • RQ4統合された知識統合戦略は、多様な視覚・言語ベンチマークにおいて一貫して性能向上をもたらすか?
  • RQ5ROSITAは、UNITERのようなベースラインモデルと比較して、アテンションベースのクロスモーダル整合性をどの程度向上させるか?

主な発見

  • ROSITAは、3つの主要な視覚・言語タスク(視覚質問応答、画像・テキスト検索、参照表現理解)における6つのベンチマークデータセットで、既存の最先端VLP手法を顕著に上回る性能を達成した。
  • 完全なROSITAモデルは、すべてのアブレーションバリアントを上回り、クロスモーダルおよびインフラモーダル知識を共同で統合する有効性を確認した。
  • 独立したマスキング確率を用いたSKM戦略は、ハイパーパrameter選択に敏感な同一確率戦略と比較して優れた性能を示した。
  • 可視化されたアテンションマップから、ROSITAは正確なクロスモーダル整合性を学習していることが示された。例えば、マスクされた' ramp '領域が正しく' ramp 'という語に一致している。一方、UNITERのようなベースラインモデルは失敗し、' skate 'などの誤ったトークンに注目している。
  • 複数の領域や語を同時にマスクした場合でも、ROSITAは正確な整合性を維持しており、' bowl & carrots 'を' bowl 'と' carrots 'に正しく関連づけ、' man ', ' candles ', ' cake 'をそれぞれの画像領域に正しく関連づけている。
Figure 2 . The flowchart of knowledge extraction given an image-text pair. It consists of two main stages, namely the unified scene graph construction and knowledge representation.
Figure 2 . The flowchart of knowledge extraction given an image-text pair. It consists of two main stages, namely the unified scene graph construction and knowledge representation.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。