[论文解读] Replication Markets: Results, Lessons, Challenges and Opportunities in AI Replication
本文提出复制市场作为一种可扩展的方法,通过利用人类对复制结果的预测来评估人工智能和机器学习研究的可重复性。采用代理评分来估计预测准确性,而无需真实结果,该方法可实现专家排名、优先开展高影响力复制工作,并通过月度奖项激励更好的可重现性,验证中观察到代理评分与真实表现之间存在强相关性。
The last decade saw the emergence of systematic large-scale replication projects in the social and behavioral sciences, (Camerer et al., 2016, 2018; Ebersole et al., 2016; Klein et al., 2014, 2018; Collaboration, 2015). These projects were driven by theoretical and conceptual concerns about a high fraction of "false positives" in the scientific publications (Ioannidis, 2005) (and a high prevalence of "questionable research practices" (Simmons, Nelson, and Simonsohn, 2011). Concerns about the credibility of research findings are not unique to the behavioral and social sciences; within Computer Science, Artificial Intelligence (AI) and Machine Learning (ML) are areas of particular concern (Lucic et al., 2018; Freire, Bonnet, and Shasha, 2012; Gundersen and Kjensmo, 2018; Henderson et al., 2018). Given the pioneering role of the behavioral and social sciences in the promotion of novel methodologies to improve the credibility of research, it is a promising approach to analyze the lessons learned from this field and adjust strategies for Computer Science, AI and ML In this paper, we review approaches used in the behavioral and social sciences and in the DARPA SCORE project. We particularly focus on the role of human forecasting of replication outcomes, and how forecasting can leverage the information gained from relatively labor and resource-intensive replications. We will discuss opportunities and challenges of using these approaches to monitor and improve the credibility of research areas in Computer Science, AI, and ML.
研究动机与目标
- 为应对人工智能和机器学习领域日益严重的不可重现结果问题,受社会科学领域‘复制危机’的启发。
- 开发可扩展的、基于市场的机制,用于预测人工智能研究主张的可重复性。
- 使用代理评分来估计预测质量,而无需访问真实结果。
- 通过月度奖项和排行榜激励高质量预测。
- 通过基于专家预测优先安排复制工作,提升研究可信度。
提出的方法
- 利用预测市场和调查收集关于人工智能/机器学习主张是否能成功复制的预测。
- 应用代理评分规则(SSR)仅使用参与者预测来估计预测准确度,而无需真实结果。
- 以Brier评分作为预测质量的基准,SSR设计为在预测量增加时收敛至真实Brier评分。
- 使用SSR评分对参与者进行排名,以识别顶尖预测者并颁发月度奖项。
- 通过将代理评分与实验数据中的真实Brier评分进行比较,验证SSR性能。
- 使用SSR加权聚合形成集体可重复性评分(CS),以反映专家共识并估计其可靠性。
实验结果
研究问题
- RQ1人类对复制结果的预测能否可靠地预测人工智能和机器学习研究的成功?
- RQ2代理评分技术能否在无真实结果访问的情况下准确估计预测质量?
- RQ3月度奖项和排行榜在维持长期预测任务参与度方面有多有效?
- RQ4基于SSR的聚合与实际复制结果的相关性如何?
- RQ5基于预测的优先排序在多大程度上能降低大规模复制工作的成本并提高其影响力?
主要发现
- 代理评分(SSR)在期望上成功恢复了Brier评分,使无需真实结果即可可靠评估预测准确度成为可能。
- 在实验数据中观察到SSR估计得分与真实Brier得分之间存在强相关性,验证了该方法的准确性。
- 月度奖项激励显著提升了参与者参与度,并维持了长期参与。
- 通过SSR识别出的顶尖预测者在参与者中持续获得认可,其成果在博客和社交媒体上获得公开认可。
- SSR加权聚合与基于市场的结果高度一致,表明其具有稳健性和信息量。
- 该方法通过识别高可信度主张以供进一步验证,实现了可扩展且成本效益高的复制工作优先排序。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。