Skip to main content
QUICK REVIEW

[論文レビュー] WenLan 2.0: Make AI Imagine via a Multimodal Foundation Model.

Nanyi Fei, Zhiwu Lu|arXiv (Cornell University)|Oct 27, 2021
Multimodal Machine Learning Applications参考文献 55被引用数 5
ひとこと要約

WenLan 2.0 は、弱い相関関係を持つウェブクロール済みの視覚的・テキスト的データ上で自己教師あり学習によって事前学習されたマルチモーダル基盤モデルを導入し、多様な認知的タスクにわたる強力なゼロショット一般化を実現する。高度な想起力と常識的推論能力を示し、人工汎用知能への重要な一歩を示している。

ABSTRACT

The fundamental goal of artificial intelligence (AI) is to mimic the core cognitive activities of human including perception, memory, and reasoning. Although tremendous success has been achieved in various AI research fields (e.g., computer vision and natural language processing), the majority of existing works only focus on acquiring single cognitive ability (e.g., image classification, reading comprehension, or visual commonsense reasoning). To overcome this limitation and take a solid step to artificial general intelligence (AGI), we develop a novel foundation model pre-trained with huge multimodal (visual and textual) data, which is able to be quickly adapted for a broad class of downstream cognitive tasks. Such a model is fundamentally different from the multimodal foundation models recently proposed in the literature that typically make strong semantic correlation assumption and expect exact alignment between image and text modalities in their pre-training data, which is often hard to satisfy in practice thus limiting their generalization abilities. To resolve this issue, we propose to pre-train our foundation model by self-supervised learning with weak semantic correlation data crawled from the Internet and show that state-of-the-art results can be obtained on a wide range of downstream tasks (both single-modal and cross-modal). Particularly, with novel model-interpretability tools developed in this work, we demonstrate that strong imagination ability (even with hints of commonsense) is now possessed by our foundation model. We believe our work makes a transformative stride towards AGI and will have broad impact on various AI+ fields (e.g., neuroscience and healthcare).

研究の動機と目的

  • 既存のマルチモーダル基盤モデルが事前学習データにおける強い意味的整合性に依存するという限界を克服すること。
  • 多様な単モーダルおよびクロスモーダル認知的タスクに一般化可能な基盤モデルを開発すること。
  • 事前学習段階で正確な画像・テキストの整合性を必要としない状況でも、AIモデルが強力な想起力と常識的推論能力を発揮できるようにすること。
  • インターネット上の弱い相関関係を持つデータに対する自己教師あり学習の有効性を示し、一般化可能なマルチモーダル表現を構築すること。
  • モデルの推論および想起能力を分析・検証できる新しい解釈可能性ツールを提供すること。

提案手法

  • インターネットからクロールした大規模で弱い相関関係を持つマルチモーダルデータを用いて、自己教師あり学習により基盤モデルを事前学習する。
  • 正確な一致を要件としないで、画像とテキスト間の弱い意味的相関を活用することで、モデルの頑健性と一般化性能を向上させる。
  • 視覚的およびテキスト的入力を統合的に符号化できるマルチモーダルアーキテクチャを設計し、多様な下流タスクをサポートする。
  • 想起および常識的推論の分野における内部表現および推論プロセスを分析するための新しい解釈可能性ツールを導入する。
  • 画像分類、視覚的質問応答、テキスト生成など、幅広い下流タスクで事前学習済みモデルを微調整する。
  • 対応するアノテーションが不要な状況でも、共同表現を学習できるように、対照的およびマスキング学習の目的関数を活用する。

実験結果

リサーチクエスチョン

  • RQ1弱い相関関係を持つウェブデータで事前学習された基盤モデルは、多様な認知的タスクで強力な性能を達成できるか?
  • RQ2このようなモデルは、明示的な教師信号なしに、どの程度の想起力と常識的推論能力を示せるか?
  • RQ3弱いアライメントデータにおける自己教師あり学習は、教師ありまたは強いアライメントを要する事前学習と比較して、一般化性能においてどのように異なるか?
  • RQ4新規に開発された解釈可能性ツールは、モデルの内部推論および想起能力を効果的に明らかに・検証できるか?
  • RQ5弱い相関関係を持つデータでの事前学習が、下流タスクにおけるゼロショットおよびフェイシュート転移性能に与える影響は何か?

主な発見

  • WenLan 2.0 モデルは、単モーダルおよびクロスモーダルの両方のベンチマークを含む、多様な下流認知的タスクで最先端の性能を達成している。
  • モデルは強力なゼロショット一般化能力を示しており、微調整なしに未学習のタスクに対しても効果的に一般化している。
  • 新規に開発された解釈可能性ツールを通じて、最小限のヒントしか与えられていない状況でも、想起力および常識的推論の兆候が観察された。
  • 弱い相関関係を持つデータに対する自己教師あり事前学習により、頑健で一般化可能なマルチモーダル表現が得られ、事前学習データで強いアライメントを要するモデルを上回った。
  • 正確な画像・テキストのアライメントがなくても、標準ベンチマークにおいて既存のモデルと同等またはそれ以上の性能を示した。
  • 解釈可能性ツールの分析から、モデルが妥当で文脈に適った視覚的およびテキスト的推論を生成できており、認知的行動に類似した挙動が顕在化していることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。