Skip to main content
QUICK REVIEW

[论文解读] Benchmarks for Automated Commonsense Reasoning: A Survey

Ernest Davis|arXiv (Cornell University)|Feb 9, 2023
Explainable Artificial Intelligence (XAI)被引用 6
一句话总结

本文综述了文本、图像、视频及模拟环境中的139项常识推理基准,识别出其在设计与评估中的关键缺陷。文章主张应开发更高质量的基准,聚焦于基础性推理能力(时间、空间、物理、心理及社会推理),而非百科知识,并呼吁实施严格的基准筛选、多语言测试及伦理数据实践,以确保对人工智能常识能力的可靠衡量。

ABSTRACT

More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems. However, these benchmarks are often flawed and many aspects of common sense remain untested. Consequently, we do not currently have any reliable way of measuring to what extent existing AI systems have achieved these abilities. This paper surveys the development and uses of AI commonsense benchmarks. We discuss the nature of common sense; the role of common sense in AI; the goals served by constructing commonsense benchmarks; and desirable features of commonsense benchmarks. We analyze the common flaws in benchmarks, and we argue that it is worthwhile to invest the work needed ensure that benchmark examples are consistently high quality. We survey the various methods of constructing commonsense benchmarks. We enumerate 139 commonsense benchmarks that have been developed: 102 text-based, 18 image-based, 12 video based, and 7 simulated physical environments. We discuss the gaps in the existing benchmarks and aspects of commonsense reasoning that are not addressed in any existing benchmark. We conclude with a number of recommendations for future development of commonsense AI benchmarks.

研究动机与目标

  • 系统性地调查并分析人工智能领域现有常识推理基准的现状。
  • 识别基准设计中的关键缺陷,这些缺陷会削弱对人工智能系统常识推理能力的可靠测量。
  • 倡导开发更高质量的基准,聚焦于基础性常识推理,而非常识或专家知识。
  • 推荐基准创建的最佳实践,包括数据筛选、多语言测试及伦理数据处理。
  • 解决基准表现与人工智能系统实际类人理解之间的错位问题。

提出的方法

  • 在四种模态中调查了139项基准:文本(102项)、图像(18项)、视频(12项)和模拟物理环境(7项)。
  • 分析现有基准中的设计缺陷,包括模糊、过于简单或谜题式的问题。
  • 评估基准的目标、理想特征及构建技术,强调一致性与代表性的重要性。
  • 提出高质量基准的标准,包括代表性采样、最小化缺陷,以及涵盖多样化的推理类型。
  • 建议基准创建者与评审者通过示例透明化、缺陷报告及样本数据集审查来实施质量控制。
  • 倡导开展多语言基准测试,以检验人工智能推理是否与语言无关,或是否依赖于语言模式。

实验结果

研究问题

  • RQ1现有常识推理基准在设计与评估方面存在哪些关键缺陷?
  • RQ2当前基准在多大程度上衡量了真正的基础性常识推理,而非表面知识或模式匹配?
  • RQ3如何系统性地提升基准质量,以确保对人工智能系统评估的可靠性与意义?
  • RQ4为何使用人类设计的谜题或测试作为人工智能系统的基准存在问题?
  • RQ5多语言评估在评估人工智能常识推理的鲁棒性与泛化能力方面应发挥何种作用?

主要发现

  • 已开发超过100项基准,但其中许多存在严重缺陷,如问题模糊、推理过于简单或具有谜题结构。
  • 大量基准依赖于广为人知的脑筋急转弯或发展心理学谜题,这些可能被大型语言模型记忆而非真正推理。
  • 许多基准未能测试基础性常识推理(如时间、空间、物理、心理及社会推理),而是聚焦于常识或百科知识。
  • 本文发现,当前基准往往不能作为实际理解的可靠代理,导致人工智能研究与评估出现误导。
  • 评审者与研究社区必须在识别和标记缺陷基准方面发挥更强作用,尤其是在这些基准被用于同行评审出版物时。
  • 当基准被用作训练数据而未提供适当的退出机制时,尤其在数据源自无偿人力劳动的情况下,会引发伦理问题。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。