[论文解读] Judging the Judges: A General Framework for Evaluating the Performance of International Sports Judges
本文提出一种通用框架,通过建模八项奥运赛事中异方差性的裁判打分误差,评估国际体育裁判的表现,揭示出裁判准确性随运动员表现水平提高而提升——但马术项目例外,其在高水平时裁判分歧反而加剧。该方法采用内在误差变异性的二次近似,并结合标准化打分,以检测异常或偏见打分,从而实现跨项目的裁判表现客观评估。
The monitoring of judges and referees in sports has become an important topic due to the increasing media exposure of international sporting events and the large monetary sums involved. In this article, we present a method to assess the accuracy of sports judges and estimate their bias. Our method is broadly applicable to all sports where panels of judges evaluate athletic performances on a finite scale. We analyze judging scores from eight different sports with comparable judging systems: diving, dressage, figure skating, freestyle skiing (aerials), freestyle snowboard (halfpipe, slopestyle), gymnastics, ski jumping and synchronized swimming. With the notable exception of dressage, we identify, for each aforementioned sport, a general and accurate pattern of the intrinsic judging error as a function of the performance level of the athlete. This intrinsic judging inaccuracy is heteroscedastic and can be approximated by a quadratic curve, indicating increased consensus among judges towards the best athletes. Using this observation, the framework developed to assess the performance of international gymnastics judges is applicable to all these sports: we can evaluate the performance of judges compared to their peers and distinguish cheating from unintentional misjudging. Our analysis also leads to valuable insights about the judging practices of the sports under consideration. In particular, it reveals a systemic judging problem in dressage, where judges disagree on what constitutes a good performance.
研究动机与目标
- 开发一种通用且可扩展的方法,用于评估多种体育项目中国际裁判的准确性。
- 量化在采用有限打分制的小组评分体系体育项目中,内在打分误差变异性的变化,作为运动员表现水平的函数。
- 通过将个体裁判的表现与同侪表现对比,利用标准化误差指标检测异常或偏见打分。
- 解决评分中的系统性问题,如国籍偏见和共识缺失,特别是在主观性较高的项目(如马术)中。
- 为体育联合会提供一种透明、数据驱动的工具,用于监控和提升裁判质量,而无需依赖匿名或不透明的系统。
提出的方法
- 将打分误差建模为异方差性随机变量,其方差取决于运动员表现水平(c)。
- 拟合一条凹二次曲线,以估计每项运动在不同表现水平下内在打分误差标准差 σd(c)。
- 使用公式 mp,j ≜ êp,j / σ̂(cp) 标准化每位裁判的误差,使其偏差相对于该表现水平下的预期变异进行缩放。
- 将裁判的打分分值 Mj 定义为标准化误差的均方根:Mj ≜ √E[mp,j²],其中 Mj = 1 表示偏离中位数一个标准差,Mj = 0 表示与同侪完全一致。
- 利用 Mj 标记异常打分(例如 >2·σ̂d(cp)·Mj),并根据裁判的固有准确性调整阈值,以区分异常行为与罕见但准确的偏见行为。
- 将估计的 σ̂d(cp) 和 Mj 整合进偏差分析中,如国籍偏见检测,以区分一致性与系统性偏袒。
实验结果
研究问题
- RQ1在不同体育项目中,打分误差的变异性如何随运动员表现水平变化?
- RQ2在采用小组评分制的体育项目中,内在打分误差变异性在多大程度上可由凹二次函数进行建模?
- RQ3为何马术项目在高水平时裁判分歧不断加剧,而其他项目则不然?
- RQ4标准化打分分值 Mj 是否能有效区分异常行为与持续偏见的裁判?
- RQ5体育联合会如何利用该框架在不依赖匿名或不透明系统的情况下,监控并提升打分准确性?
主要发现
- 在八项奥运赛事中(马术除外),随着运动员表现水平提高,打分误差的变异性降低,且符合凹二次曲线,表明高水平运动员的裁判共识度更高。
- 马术是显著例外:在高水平时,裁判分歧持续加剧,表明其主观性强,且对高质量表现的评判缺乏共识。
- 内在打分误差标准差 σd(c) 可通过二次模型准确近似,从而可靠估计各表现水平下的预期误差变异性。
- 打分分值 Mj 有效量化了裁判的准确性:Mj = 1 表示偏离中位数一个标准误差,Mj = 0 表示与同侪完全一致。
- 可通过裁判的 Mj 调整异常值检测阈值,从而识别极端误判,同时区分异常行为与罕见但准确的偏见判断。
- 该框架实现了对裁判的客观、透明且可扩展的监控,为检测小组评分制体育项目中的能力不足或潜在腐败行为提供了有效工具。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。