Skip to main content
QUICK REVIEW

[論文レビュー] Language Is Not All You Need: Aligning Perception with Language Models

Shaohan Huang, Dong Li|arXiv (Cornell University)|Feb 27, 2023
Multimodal Machine Learning Applications被引用数 164
ひとこと要約

Kosmos-1 は、ウェブ規模のテキスト、画像、およびインタリーブされたマルチモーダルデータを用いてゼロから訓練されたマルチモーダル大規模言語モデルであり、ファインチューニングなしで言語、知覚、ビジョンを横断するゼロショットおよび少数ショット推論を行う。

ABSTRACT

A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce Kosmos-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train Kosmos-1 from scratch on web-scale multimodal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that Kosmos-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs.

研究の動機と目的

  • マルチモーダル知覚と言語モデルを整合させるモデルの必要性を動機づけ、人工知能全般知性を追求する。
  • 一般的なモダリティを認識し、指示に従い、ファインチューニングなしでコンテキスト内で学習できるモデルを開発する。
  • 言語モデルがモダリティを横断する普遍的なタスクインタフェースとして機能できることを示す。
  • 言語のみとマルチモーダル機能間のクロスモーダル転移の利点を示す。
  • MLLMs のためのベンチマークと新しい Raven IQ 風の非言語推論データセットを提供する。

提案手法

  • Kosmos-1 を、テキストと画像が交互に含まれるデータ、画像キャプションの対、テキストのみデータを含むウェブ規模のマルチモーダルコーパスでゼロから訓練する。
  • 埋め込まれたマルチモーダル入力を持つ Transformer ベースの因果言語モデルを中核インタフェースとして用いる。
  • 安定性と長い文脈のモデリングを改善するために Magneto バックボーンと xPos 相対位置エンコーディングを採用する。
  • 混在モダリティ上で次トークン予測の事前学習を行い、学習のために離散トークン損失を維持する。
  • 言語のみの命令チューニングを実施して指示追従を改善し、マルチモーダルタスクへの転移を促す。

実験結果

リサーチクエスチョン

  • RQ1マルチモーダル大規模言語モデル(MLLM)は、知覚を言語モデルと整合させて、ファインチューニングなしで言語タスクとビジョンタスクの両方を実行できるか。
  • RQ2クロスモーダル転移が、言語・知覚言語タスクをどの程度改善し得るか、そしてその逆もどの程度可能か。
  • RQ3MLLM は、非言語推論、OCRなしタスク、知覚ベースの推論を、テキストのみのモデルと比べてどの程度上手く扱えるか。

主な発見

  • Kosmos-1 は、勾配更新なしで、言語、知覚言語、ビジョンタスク全般におけるゼロショットおよび少数ショット能力を示す。
  • モデルはクロスモーダル転移の恩恵を受け、言語機能とマルチモーダル機能が相互に補完し合う。
  • A Raven IQ-style 非言語推論ベンチマークは Kosmos-1 がゼロショットの非言語推論を行えることを示し、視覚-テキスト文脈での抽象的パターン認識を示唆する。
  • OCR-free タスク such as rendered text and web-page understanding are feasible with Kosmos-1, without external tools.
  • Multimodal chain-of-thought prompting improves performance on perception-language tasks by generating intermediate rationale before final answers.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。