[论文解读] The Limitations of Standardized Science Tests as Benchmarks for Artificial Intelligence Research: Position Paper
本文认为,像SAT或州立考试这样的标准化科学测试作为衡量人工智能科学理解能力进展的基准并不理想,因为它们强调的是人类的困难,而非人工智能所缺乏的根本性世界知识。文章主张采用更丰富、更多元化的基准,如文本理解、物理推理任务和机器人整合,而非考试形式的问题,以推动人工智能在科学推理方面实现更有意义的进展。
In this position paper, I argue that standardized tests for elementary science such as SAT or Regents tests are not very good benchmarks for measuring the progress of artificial intelligence systems in understanding basic science. The primary problem is that these tests are designed to test aspects of knowledge and ability that are challenging for people; the aspects that are challenging for AI systems are very different. In particular, standardized tests do not test knowledge that is obvious for people; none of this knowledge can be assumed in AI systems. Individual standardized tests also have specific features that are not necessarily appropriate for an AI benchmark. I analyze the Physics subject SAT in some detail and the New York State Regents Science test more briefly. I also argue that the apparent advantages offered by using standardized tests are mostly either minor or illusory. The one major real advantage is that the significance is easily explained to the public; but I argue that even this is a somewhat mixed blessing. I conclude by arguing that, first, more appropriate collections of exam style problems could be assembled, and second, that there are better kinds of benchmarks than exam-style problems. In an appendix I present a collection of sample exam-style problems that test kinds of knowledge missing from the standardized tests.
研究动机与目标
- 挑战标准化科学测试是衡量人工智能在科学理解方面进展的有效基准这一假设。
- 指出人工智能系统往往缺乏‘任何傻子都懂’的常识性知识,而擅长形式化科学,导致基准的相关性不匹配。
- 主张公众基于测试表现对人工智能成功的认知是误导性的,可能对领域信誉造成损害。
- 提出标准化测试之外的更好替代方案,包括多样化的推理任务和现实世界应用。
- 倡导在人工智能研究中摒弃标准化测试的限制,以实现更具创新性和意义的进展。
提出的方法
- 分析物理SAT和纽约州立科学考试,识别其中对基本世界知识的遗漏。
- 指出标准化测试排除了‘显而易见’的物理常识(例如:你无法把一个西瓜塞进一个三明治袋中),而这些常识对人类般的理解至关重要。
- 提出一组经筛选的、以考试形式呈现的问题,用于测试这些缺失的常识性知识,详见附录A。
- 倡导采用超越选择题考试的基准,如论述题、文本理解,以及与规划系统或设计系统的整合。
- 强调需要超越表面的测试表现,深入到深层推理、知识表征和与现实物理世界的互动。
- 建议人工智能研究人员避免对官方测试签署保密协议,以确保评估的透明度和可复现性。
实验结果
研究问题
- RQ1为何像SAT或州立考试这样的标准化科学测试,作为评估人工智能系统科学推理能力的基准是不充分的?
- RQ2哪些根本性的物理知识——即‘任何傻子都懂’的知识——在标准化测试中系统性地被遗漏,但对人类般的理解至关重要?
- RQ3公众对人工智能表现的认知,受其通过测试的宣传影响,如何扭曲了人工智能真实进展的状况?
- RQ4哪些替代性基准能更好地衡量并推动人工智能在科学推理和世界知识方面的进步?
- RQ5为何人工智能研究采用标准化测试的限制(如固定格式和保密协议)是适得其反的?
主要发现
- 标准化科学测试与人工智能能力不匹配,因为它们强调的是人类的困难,而非人工智能通常缺乏的常识性知识。
- 人工智能系统可以掌握形式化的科学方程,但在涉及日常物品和因果关系的基本物理推理方面可能失败,例如无法理解把一个西瓜塞进三明治袋是不现实的。
- 公众常常将人工智能通过标准化测试误解为具备人类水平智能的证据,导致误导性新闻标题和过高的期望。
- 官方测试的保密协议会阻碍透明度并妨碍可复现性,使这些测试不适合开放的人工智能研究。
- 标准化测试最大的好处——公众认可——被误报的风险和研究自由的丧失所超过。
- 更有效的基准应包括文本理解、基于论文的推理、物理情境的多样化变化,以及与规划和机器人系统集成。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。