Skip to main content
QUICK REVIEW

[論文レビュー] On the Measure of Intelligence

François Chollet|arXiv (Cornell University)|Nov 5, 2019
Computability, Logic, AI Algorithms参考文献 48被引用数 10
ひとこと要約

本稿では、アルゴリズム的情報理論に基づき、技能習得効率としての知能の形式的定義を提示し、人間らしく一般化する知能を測定するためのベンチマークとして、抽象化と推論コーパス(ARC)を導入する。ARCは、事前知識、範囲、一般化の難易度を標準化することで、タスク特有のスキルにとどまらない広範な認知的能力を評価可能な公平な比較を可能にする。

ABSTRACT

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans. Over the past hundred years, there has been an abundance of attempts to define and measure intelligence, across both the fields of psychology and AI. We summarize and critically assess these definitions and evaluation approaches, while making apparent the two historical conceptions of intelligence that have implicitly guided them. We note that in practice, the contemporary AI community still gravitates towards benchmarking intelligence by comparing the skill exhibited by AIs and humans at specific tasks such as board games and video games. We argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to "buy" arbitrary levels of skills for a system, in a way that masks the system's own generalization power. We then articulate a new formal definition of intelligence based on Algorithmic Information Theory, describing intelligence as skill-acquisition efficiency and highlighting the concepts of scope, generalization difficulty, priors, and experience. Using this definition, we propose a set of guidelines for what a general AI benchmark should look like. Finally, we present a benchmark closely following these guidelines, the Abstraction and Reasoning Corpus (ARC), built upon an explicit set of priors designed to be as close as possible to innate human priors. We argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans.

研究の動機と目的

  • AI研究における、厳密で実行可能かつ定量的な知能の定義が不足している問題に取り組むこと。
  • 無制限のデータと事前知識に依存して人工的にスキルが高められるタスク特有のスキルにのみ焦点を当てる現在のベンチマーク手法の限界を克服すること。
  • コア知識の事前知識と一般化の難易度を標準化することで、AIと人間の知能を公平に比較できるベンチマークを開発すること。
  • 特定のタスクでの高いスキルを達成することから、人間らしく広範な一般化と抽象化能力を持つシステムの開発へ研究の焦点を転換すること。
  • アルゴリズム的情報理論に基づく形式的で定量的な枠組みを提供し、一般AIの開発を導くこと。

提案手法

  • 知能を、経験と事前知識から新しいスキルを習得する速度として定義する。
  • 学習パフォーマンスを包括的に評価するため、計算、時間、エネルギー、リスクの4つの効率次元を導入する。
  • 最小プログラム長と圧縮不能な複雑性の観点から知能を定量的に測定するため、アルゴリズム的情報理論を用いた形式的フレームワークを構築する。
  • 共通のインナーヒューリスティックなコア知識事前知識を仮定する、新規一般化を要請するタスクを備えたベンチマークとして、抽象化と推論コーパス(ARC)を設計する。
  • 記憶ではなく一般化をテストする、新規性と難易度に配慮したタスク生成を可能にする、教師-生徒学習ループを用いる。
  • すべての受験者が同一の基礎的知識を前提とするように、範囲、事前知識、経験を制御することで、AIシステムと人間の間の公平な比較を確保する。

実験結果

リサーチクエスチョン

  • RQ1知能を、単なる記述的ではなく、実行可能で測定可能かつ説明可能な方法で形式的に定義するにはどうすればよいか?
  • RQ2なぜタスク特有のスキルの測定では一般知能の評価が不十分であり、現在のベンチマークパラダイムの限界は何か?
  • RQ3公平で定量的な知能評価において、事前知識、経験、一般化の難易度を体系的にどのように扱えるか?
  • RQ4抽象化と推論コーパス(ARC)が、AIシステムにおける人間らしく一般化された流動的知能を測定する有効なベンチマークとしてどの程度機能するか?
  • RQ5コア知識事前知識が、人間と人工知能の間で公平で意味のある比較を可能にする役割を果たすメカニズムは何か?

主な発見

  • 本稿では、真の一般化能力を反映させるため、知能は生のスキルではなく、スキル習得効率として測定すべきであると確立する。
  • 現在のAIベンチマークは、無制限のデータと事前知識を許容することで人工的にスキルが高められ、一般化の限界が隠蔽されるため、タスク特有のパフォーマンスに依存するのでは不十分である。
  • ARCベンチマークは、人間には解けるが、現在の機械学習モデル(深層学習を含む)には解けないという特徴を持ち、広範な一般化能力のギャップを示している。
  • ARCは、共通のインナーヒューリスティックな人間らしさの事前知識(コア知識)を強制することで、AIシステムと人間の間で公平な比較を可能にする。
  • アルゴリズム的情報理論に基づく形式的フレームワークは、知能と一般化効率についての厳密な定量的根拠を提供する。
  • 新規性と難易度に配慮したタスク生成を可能にする教師-生徒ループは、カリキュラム学習と一般知能のベンチマークに向けたスケーラブルで開放的なアプローチを提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。