Skip to main content
QUICK REVIEW

[论文解读] HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning

Oguzhan Gencoglu, Mark van Gils|arXiv (Cornell University)|Apr 16, 2019
Explainable Artificial Intelligence (XAI)参考文献 42被引用 18
一句话总结

本文识别出HARKing(结果已知后才提出假设)是深度学习研究中普遍存在的问题,其根源在于竞争压力、发表压力以及有缺陷的基准测试。文章倡导推动文化变革,强调事前假设制定、可复现性以及伦理AI实践,以提升机器学习系统的科学严谨性与可信度。

ABSTRACT

Recent advancements in machine learning research, i.e., deep learning, introduced methods that excel conventional algorithms as well as humans in several complex tasks, ranging from detection of objects in images and speech recognition to playing difficult strategic games. However, the current methodology of machine learning research and consequently, implementations of the real-world applications of such algorithms, seems to have a recurring HARKing (Hypothesizing After the Results are Known) issue. In this work, we elaborate on the algorithmic, economic and social reasons and consequences of this phenomenon. We present examples from current common practices of conducting machine learning research (e.g. avoidance of reporting negative results) and failure of generalization ability of the proposed algorithms and datasets in actual real-life usage. Furthermore, a potential future trajectory of machine learning research and development from the perspective of accountable, unbiased, ethical and privacy-aware algorithmic decision making is discussed. We would like to emphasize that with this discussion we neither claim to provide an exhaustive argumentation nor blame any specific institution or individual on the raised issues. This is simply a discussion put forth by us, insiders of the machine learning field, reflecting on us.

研究动机与目标

  • 识别并分析HARKing在深度学习研究中的普遍性,特别是在高风险应用场景中的表现。
  • 考察超竞争的研究环境与发表激励如何导致事后假设形成及负面结果的隐瞒。
  • 指出当前模型与数据集在现实场景中缺乏泛化能力,从而削弱科学有效性。
  • 倡导在机器学习研究方法的早期阶段整合伦理、可解释性及隐私保护原则。
  • 推动机器学习研究在文化和制度层面的改革,以增强问责性、透明度与科学诚信。

提出的方法

  • 分析深度学习研究中的常见实践,包括对正面结果的选择性报告以及对负面结果的未发表现象。
  • 探讨自动化机器学习(AutoML)在通过自动超参数搜索与模型选择可能加剧HARKing问题中的作用。
  • 回顾单一指标基准测试的局限性,以及因数据集与评估偏差导致的模型性能误报。
  • 提出将事前假设形成整合进研究生命周期,以契合假说-演绎科学方法。
  • 讨论可复现性、可解释性以及隐私保护技术(如联邦学习、差分隐私)作为防范HARKing的保障措施的重要性。
  • 评估政策机制(如GDPR的解释权与可操作审计)在提升AI系统问责性方面的有效性。

实验结果

研究问题

  • RQ1HARKing在深度学习研究中在多大程度上存在?其在研究文化与激励机制中的根本原因是什么?
  • RQ2发表压力与对最先进性能的追求如何导致负面结果被隐瞒及事后假设的生成?
  • RQ3为何当前的深度学习模型与数据集常常无法泛化到现实应用场景?这与HARKing有何关联?
  • RQ4自动化机器学习(AutoML)在哪些方面既能缓解又能加剧HARKing行为?
  • RQ5如何将伦理、可解释性及隐私感知的AI原则嵌入研究过程,以减少HARKing并提升科学诚信?

主要发现

  • 由于激烈的竞争、发表偏倚以及在基准上取得最先进结果的压力,HARKing在深度学习研究中普遍存在。
  • 对负面结果的隐瞒与对性能提升的片面报告严重损害可复现性,并扭曲科学记录。
  • 当前的数据集与模型评估常常无法泛化到现实场景,表明基准性能与实际应用价值之间存在显著差距。
  • 自动化机器学习(AutoML)可减少人工调参,但如果未与事前假设检验及透明评估相结合,也可能助长HARKing。
  • 可解释性与问责机制(如GDPR的解释权、可操作审计)有助于检测与缓解HARKing,但需稳健实施。
  • 将联邦学习与差分隐私等隐私保护技术整合进研究方法,可增强科学诚信,并降低数据滥用风险。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。