[论文解读] AI Evaluation: past, present and future
本文批判了传统的任务导向型人工智能评估方法,主张转向能力导向的评估,强调系统性、稳健性与非人类中心的测量方式。文章基于心理测量学原理、算法信息论与标准化基准,提出了一套基于认知能力的AI系统评估框架,旨在推动窄域与通用人工智能的发展。
Artificial intelligence develops techniques and systems whose performance must be evaluated on a regular basis in order to certify and foster progress in the discipline. We will describe and critically assess the different ways AI systems are evaluated. We first focus on the traditional task-oriented evaluation approach. We see that black-box (behavioural evaluation) is becoming more and more common, as AI systems are becoming more complex and unpredictable. We identify three kinds of evaluation: Human discrimination, problem benchmarks and peer confrontation. We describe the limitations of the many evaluation settings and competitions in these three categories and propose several ideas for a more systematic and robust evaluation. We then focus on a less customary (and challenging) ability-oriented evaluation approach, where a system is characterised by its (cognitive) abilities, rather than by the tasks it is designed to solve. We discuss several possibilities: the adaptation of cognitive tests used for humans and animals, the development of tests derived from algorithmic information theory or more general approaches under the perspective of universal psychometrics.
研究动机与目标
- 批判性评估现有的AI评估方法,特别是基于任务的评估方式,如基准测试与竞赛。
- 识别当前评估实践的局限性,包括抽样偏差、缺乏标准化以及过度依赖人类判断。
- 倡导一种系统性、计算性且非人类中心的AI评估方法,以支持窄域与通用AI的发展。
- 提出一种新范式——能力导向评估,基于认知测试、通用心理测量学与算法信息论。
- 制定设计更可靠、可复现、透明的AI系统评估框架的指导原则。
提出的方法
- 将AI评估分为三类:人类判别、问题基准与同行对抗,分析其优缺点。
- 提出一个正式的评估框架,包含问题集 $\Omega$、系统 $M$、响应函数 $\Phi$ 与性能度量 $p$,其中 $R$ 为评估情境。
- 建议使用内在难度函数与项目反应理论,建模系统在不同问题难度下的表现。
- 推荐高效的采样策略(如聚类或范围采样)以提升评估效率与稳健性。
- 主张公开评估流程与问题(但不包括完整的问题集或参数),以支持纵向比较。
- 强调通过项目反应函数与代理响应函数进行事后分析,以检测异常并验证结果。
实验结果
研究问题
- RQ1如何使AI评估更具系统性、可靠性,并减少对临时基准的依赖?
- RQ2当前基于任务的评估方法(如竞赛与基准测试)在衡量真正智能方面存在哪些局限性?
- RQ3如何形式化能力导向的评估,以评估AI系统的通用认知能力?
- RQ4通用心理测量学与算法信息论在设计客观、领域无关的AI评估方法中可发挥何种作用?
- RQ5评估框架如何确保AI系统性能的可复现性、透明性与长期可比性?
主要发现
- 传统的任务导向型评估方法(如竞赛与基准测试)虽广泛使用,但存在抽样偏差、缺乏标准化及泛化能力有限等问题。
- 人类判别评估具有主观性,易受偏见影响,尤其在评估复杂或新颖的AI行为时。
- 同行对抗评估(如AI竞赛)虽有效,但需精心设计匹配调度与自适应机制以确保稳健性。
- 所提出的基于 $\Omega$、$M$、$\Phi$、$p$ 与 $R$ 的框架,可实现形式化、可重复且可分析的评估过程,支持实证分析与比较。
- 可基于实证结果构建项目反应函数与代理响应函数,以检测异常并验证评估过程。
- 公开详细的评估结果(超越聚合得分)对于透明度、可复现性与社区验证至关重要,符合开放科学原则。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。