Skip to main content
QUICK REVIEW

[Paper Review] AI Oversight and Human Mistakes: Evidence from Centre Court

David Almog, Romain Gauriot|arXiv (Cornell University)|Jan 30, 2024
Law, Economics, and Judicial Systems5 citations
TL;DR

The paper analyzes how Hawk-Eye AI oversight in professional tennis affect umpire decision-making, showing overall mistake reductions but a shift in error types driven by psychological costs of being overruled by AI.

ABSTRACT

Powered by the increasing predictive capabilities of machine learning algorithms, artificial intelligence (AI) systems have the potential to overrule human mistakes in many settings. We provide the first field evidence that the use of AI oversight can impact human decision-making. We investigate one of the highest visibility settings where AI oversight has occurred: Hawk-Eye review of umpires in top tennis tournaments. We find that umpires lowered their overall mistake rate after the introduction of Hawk-Eye review, but also that umpires increased the rate at which they called balls in, producing a shift from making Type II errors (calling a ball out when in) to Type I errors (calling a ball in when out). We structurally estimate the psychological costs of being overruled by AI using a model of attention-constrained umpires, and our results suggest that because of these costs, umpires cared 37% more about Type II errors under AI oversight.

Motivation & Objective

  • Motivate understanding of how AI oversight influences human decision-making in high-stakes settings.
  • Quantify the impact of Hawk-Eye on umpire mistake rates and error types (serves vs. non-serves, close calls).
  • Develop a rational inattention model capturing psychological costs of being overruled by AI.
  • Assess heterogeneity by player stature and tournament stage.
  • Provide policy implications on AI oversight design and incentive alignment.

Proposed method

  • Use a two-period setting before/after Hawk-Eye introduction to identify AI oversight effects.
  • Consolidate three data sources (Hawk-Eye Base, Challenge data, video-audited merging) for point-level decisions.
  • Estimate OLS models of incorrect calls with distance bins, speed, score, and match controls to capture PostHK effects.
  • Separate analysis for serves vs. non-serves to understand task-specific effects.
  • Structurally estimate psychological costs of AI oversight via a rational-inattention model with asymmetric attention costs and AI-overrule penalties.
Figure 1 : Incorrect call rates by proximity to the line. Each dot is the rate of incorrect calls for a bin of 20 mm. Dots to the left of the dashed line represent bins out of bounds, and the right of the dashed line represents bins in bounds.
Figure 1 : Incorrect call rates by proximity to the line. Each dot is the rate of incorrect calls for a bin of 20 mm. Dots to the left of the dashed line represent bins out of bounds, and the right of the dashed line represents bins in bounds.

Experimental results

Research questions

  • RQ1Does Hawk-Eye AI oversight reduce umpire mistake rates overall?
  • RQ2How does AI oversight affect error types for close calls (serves vs. non-serves) and distance to line?
  • RQ3Do psychological costs of being overruled by AI explain changes in umpire behavior?
  • RQ4Is there heterogeneity in AI oversight effects by player ranking or tournament stage?

Key findings

  • Umpires reduce overall mistake rate by 8% (1.1 percentage points) after Hawk-Eye introduction, consistent with rational inattention.
  • For the closest calls (within 20 mm), the mistake rate increases by 22.9% (7.3 percentage points) under AI oversight.
  • Umpires increase the rate of calling balls in for close calls by 12.6% (6.2 percentage points) after Hawk-Eye, shifting errors from Type II to Type I.
  • Serves show no significant performance impact from AI oversight in the primary specification, but non-serves see a 2.3 percentage point decrease in incorrect calls (about 17% reduction from baseline).
  • Structural estimates suggest psychological costs make umpires care twice as much about Type II errors after AI oversight.
(a) Serves.
(a) Serves.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.