Skip to main content
QUICK REVIEW

[论文解读] On the Measure of Intelligence

François Chollet|arXiv (Cornell University)|Nov 5, 2019
Computability, Logic, AI Algorithms参考文献 48被引用 10
一句话总结

本文基於算法信息论,提出了一種智力的形式化定義:技能獲取效率,並引入抽象與推理語料庫(ARC)作為衡量類人通用智力的基準。ARC 透過標準化先驗知識、範圍與泛化難度,確保公平比較,使評估超越特定任務技能,涵蓋廣泛認知能力。

ABSTRACT

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans. Over the past hundred years, there has been an abundance of attempts to define and measure intelligence, across both the fields of psychology and AI. We summarize and critically assess these definitions and evaluation approaches, while making apparent the two historical conceptions of intelligence that have implicitly guided them. We note that in practice, the contemporary AI community still gravitates towards benchmarking intelligence by comparing the skill exhibited by AIs and humans at specific tasks such as board games and video games. We argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to "buy" arbitrary levels of skills for a system, in a way that masks the system's own generalization power. We then articulate a new formal definition of intelligence based on Algorithmic Information Theory, describing intelligence as skill-acquisition efficiency and highlighting the concepts of scope, generalization difficulty, priors, and experience. Using this definition, we propose a set of guidelines for what a general AI benchmark should look like. Finally, we present a benchmark closely following these guidelines, the Abstraction and Reasoning Corpus (ARC), built upon an explicit set of priors designed to be as close as possible to innate human priors. We argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans.

研究动机与目标

  • 解決人工智能研究中缺乏嚴謹、可操作且可量化的智力定義問題。
  • 克服當前基準評估僅聚焦於特定任務技能的局限,這些技能可能因無限數據與先驗知識而人為膨脹,從而掩蓋泛化能力的限制。
  • 開發一個基準,透過標準化核心知識先驗與泛化難度,實現對人工智慧與人類智力的公平比較。
  • 引導研究焦點從追求在狹窄任務中達成高技能,轉向發展具備類人廣泛泛化與抽象能力的系統。
  • 提供基於算法信息論的正式、可量化的框架,以指導通用人工智慧的發展。

提出的方法

  • 將智力定義為技能獲取效率,即系統從經驗與先驗知識中獲取新技能的速率。
  • 引入四種效率維度:計算效率、時間效率、能源效率與風險效率,以全面評估學習表現。
  • 建立基於算法信息論的形式化框架,以最小程式長度與不可壓縮複雜性來量化智力。
  • 設計抽象與推理語料庫(ARC)作為基準,其任務需具備新穎的泛化能力,並假設共享一組內在的人類類似核心知識先驗。
  • 使用師生學習循環生成具有新穎性與難度感知的創新、具挑戰性且可解的任務,專注於測試泛化能力而非記憶能力。
  • 控制範圍、先驗知識與經驗,以確保人工智慧系統與人類之間的公平比較,所有測試參與者均假設具備相同的基礎知識。

实验结果

研究问题

  • RQ1如何形式化定義智力,使其具備可操作性、可測量性與解釋性,而非僅僅是描述性?
  • RQ2為何僅測量特定任務技能不足以評估通用智力?當前基準評估模式存在哪些局限?
  • RQ3如何系統性地在公平且可量化的基礎上,考慮先驗知識、經驗與泛化難度,以評估智力?
  • RQ4抽象與推理語料庫(ARC)在多大程度上可作為衡量人工智慧系統類人通用流體智力的有效基準?
  • RQ5核心知識先驗在實現人類與人工智慧之間公平且有意義的比較中發揮何種作用?

主要发现

  • 本文確立智力應以技能獲取效率來衡量,而非原始技能,以反映真正的泛化能力。
  • 當前依賴特定任務表現的AI基準不夠充分,因其允許無限數據與先驗知識人為膨脹技能,從而掩蓋泛化能力的極限。
  • ARC基準設計為人類可解,但對當前機器學習模型(包括深度學習)而言仍屬難解,顯示在廣泛泛化能力上存在差距。
  • ARC強制採用一組內在的人類類似先驗(核心知識),使人工智慧系統與人類在公平基礎上進行比較。
  • 基於算法信息論的形式化框架為智力與泛化效率的推理提供了嚴謹且可量化的基礎。
  • 具備新穎性與難度感知的師生循環任務生成方法,提供了一種可擴展、開放式的課程學習與通用智力基準評估方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。