[论文解读] AI Oversight and Human Mistakes: Evidence from Centre Court
本文分析职业网球中 Hawk-Eye AI 监督如何影响裁判决策,显示总体错误率下降,但错误类型发生转变,源于被 AI 推翻的心理成本。
Powered by the increasing predictive capabilities of machine learning algorithms, artificial intelligence (AI) systems have the potential to overrule human mistakes in many settings. We provide the first field evidence that the use of AI oversight can impact human decision-making. We investigate one of the highest visibility settings where AI oversight has occurred: Hawk-Eye review of umpires in top tennis tournaments. We find that umpires lowered their overall mistake rate after the introduction of Hawk-Eye review, but also that umpires increased the rate at which they called balls in, producing a shift from making Type II errors (calling a ball out when in) to Type I errors (calling a ball in when out). We structurally estimate the psychological costs of being overruled by AI using a model of attention-constrained umpires, and our results suggest that because of these costs, umpires cared 37% more about Type II errors under AI oversight.
研究动机与目标
- 激发对 AI 监督如何影响高风险情境下人类决策的理解。
- 量化 Hawk-Eye 对裁判错误率和错误类型(发球 vs. 非发球、近拍判决)的影响。
- 建立一个捕捉被 AI 推翻的心理成本的理性注意模型。
- 评估按球员等级/身材与锦标赛阶段的异质性。
- 就AI监督设计和激励对齐提供政策含义。
提出的方法
- 使用 Hawk-Eye 引入前后两个时间段的设定来识别 AI 监督效应。
- 整合三种数据源(Hawk-Eye Base、挑战数据、视频审核合并)用于单点决策。
- 估计带有距离分箱、速度、比分和比赛控制变量的错误判决 OLS 模型,以捕捉 PostHK 效应。
- 对发球与非发球分开分析,以理解任务特定效应。
- 通过带有不对称注意成本和 AI 推翻惩罚的理性注意模型,对 AI 监督的心理成本进行结构性估计。

实验结果
研究问题
- RQ1Hawk-Eye AI 监督是否会整体降低裁判错误率?
- RQ2AI 监督如何影响近拍判决的错误类型(发球 vs. 非发球)以及与线的距离?
- RQ3被 AI 推翻的心理成本是否能解释裁判行为的变化?
- RQ4AI 监督效应在球员排名或锦标赛阶段是否存在异质性?
主要发现
- 裁判在引入 Hawk-Eye 后总体错误率下降 8%(1.1 个百分点),与理性注意一致。
- 对于最接近的判决(在 20 毫米内),在 AI 监督下错误率上升 22.9%(7.3 个百分点)。
- 在 Hawk-Eye 之后,裁判将接近判决判定为界内的球的概率上升 12.6%(6.2 个百分点),将错误从 Type II 转移到 Type I。
- 在主要估计中,发球对 AI 监督的性能影响不显著,但非发球的错误判决降低了 2.3 个百分点(相当于基线约 17% 的减小)。
- 结构性估计表明,心理成本使裁判在 AI 监督后对 Type II 错误的在意程度增至原来的两倍。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。