Skip to main content
QUICK REVIEW

[论文解读] NIPS - Not Even Wrong? A Systematic Review of Empirically Complete Demonstrations of Algorithmic Effectiveness in the Machine Learning and Artificial Intelligence Literature

Franz J. Király, Bilal A. Mateen|arXiv (Cornell University)|Dec 18, 2018
Explainable Artificial Intelligence (XAI)参考文献 76被引用 5
一句话总结

本篇系统性综述评估了2017年NeurIPS会议中121篇监督学习论文在算法有效性主张上的实证严谨性。研究发现,仅有2篇论文(1.6%)提供了完整的论证链条来支持其优于最先进基线的主张,揭示了顶级机器学习/人工智能研究在报告标准和实证验证方面存在广泛缺陷。

ABSTRACT

Objective: To determine the completeness of argumentative steps necessary to conclude effectiveness of an algorithm in a sample of current ML/AI supervised learning literature. Data Sources: Papers published in the Neural Information Processing Systems (NeurIPS, née NIPS) journal where the official record showed a 2017 year of publication. Eligibility Criteria: Studies reporting a (semi-)supervised model, or pre-processing fused with (semi-)supervised models for tabular data. Study Appraisal: Three reviewers applied the assessment criteria to determine argumentative completeness. The criteria were split into three groups, including: experiments (e.g real and/or synthetic data), baselines (e.g uninformed and/or state-of-art) and quantitative comparison (e.g. performance quantifiers with confidence intervals and formal comparison of the algorithm against baselines). Results: Of the 121 eligible manuscripts (from the sample of 679 abstracts), 99\% used real-world data and 29\% used synthetic data. 91\% of manuscripts did not report an uninformed baseline and 55\% reported a state-of-art baseline. 32\% reported confidence intervals for performance but none provided references or exposition for how these were calculated. 3\% reported formal comparisons. Limitations: The use of one journal as the primary information source may not be representative of all ML/AI literature. However, the NeurIPS conference is recognised to be amongst the top tier concerning ML/AI studies, so it is reasonable to consider its corpus to be representative of high-quality research. Conclusion: Using the 2017 sample of the NeurIPS supervised learning corpus as an indicator for the quality and trustworthiness of current ML/AI research, it appears that complete argumentative chains in demonstrations of algorithmic effectiveness are rare.

研究动机与目标

  • 评估得出某项新机器学习/人工智能算法有效的论证链条的完整性。
  • 评估顶级机器学习会议中的监督学习论文是否提供了充分的实证证据来支持其算法优越性的主张。
  • 识别在机器学习/人工智能研究中基线报告、性能度量和统计比较方面缺失的组成部分。
  • 考察当前机器学习/人工智能出版实践是否助长了不完整或误导性的有效性主张。
  • 倡导改进报告标准,以增强机器学习/人工智能研究的可信度和可复现性。

提出的方法

  • 系统筛选了NeurIPS 2017年会议的679篇论文摘要,以识别符合条件的监督学习研究。
  • 选取了121篇报告了(半)监督模型或与这类模型融合的预处理方法用于表格数据的论文。
  • 采用三级评估框架:实验(真实/合成数据)、基线(无信息基线/最先进基线)以及定量比较(带置信区间的性能度量与正式检验)。
  • 三位独立评审员根据预设标准评估每篇论文的论证完整性。
  • 在筛选和全文评估过程中,通过共识机制和第三方仲裁解决分歧。
  • 在评审后进行统计分析,量化整个文献集合中的报告缺陷。

实验结果

研究问题

  • RQ1NeurIPS 2017年监督学习论文在多大程度上提供了支持算法有效性主张的完整实证论证?
  • RQ2在顶级机器学习论文中,无信息基线、最先进基线、置信区间和正式统计比较的报告频率如何?
  • RQ3NeurIPS 2017年文献集中有多少比例的论文提出了可完全测试且科学有效的算法优越性论证?
  • RQ4尽管已有既定的评估标准,为何大多数机器学习/人工智能论文仍未能提供完整的论证链条?
  • RQ5出版与评审实践中存在哪些系统性问题,导致机器学习/人工智能研究中‘甚至不正确’的主张普遍存在?

主要发现

  • 在121篇符合条件的论文中,仅有2篇(1.6%)提供了完整论证链条,以支持其新算法有效的主张。
  • 99%的论文使用了真实世界数据,而仅有29%使用了合成数据,表明对真实数据的强烈依赖,缺乏系统的消融分析。
  • 91%的论文未报告无信息基线,55%报告了最先进基线,表明比较严谨性较弱。
  • 仅32%的论文报告了性能度量的置信区间,且无一提供其计算的合理性说明或参考文献。
  • 仅3%的论文进行了正式的统计比较,以验证与基线的性能差异。
  • 本研究结论指出,在高影响力机器学习/人工智能研究中,算法有效性主张的完整实证论证极为罕见,大多数主张因缺乏基础证据而属于‘甚至不正确’的范畴。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。