Skip to main content
QUICK REVIEW

[论文解读] Judging the Judges: Evaluating the Performance of International Gymnastics Judges

Hugues Mercier, Sandro Heiniger|arXiv (Cornell University)|Jul 26, 2018
Sports Analytics and Performance参考文献 29被引用 9
一句话总结

本文介紹了裁判評估計劃(JEP),這是一套統計引擎,利用基於異方差隨機變數的評分機制,評估國際體操裁判的準確性、檢測偏見並確保公平性。與先前系統相比,本研究透過建模裁判內在評分誤差的變異性,實現客觀、長期的表現追蹤,進而提升國際體操聯合會(FIG)在選拔裁判與規則修訂方面的決策品質。

ABSTRACT

Judging a gymnastics routine is a noisy process, and the performance of judges varies widely. In collaboration with the Fédération Internationale de Gymnastique (FIG) and Longines, we are designing and implementing an improved statistical engine to analyze the performance of gymnastics judges during and after major competitions like the Olympic Games and the World Championships. The engine, called the Judge Evaluation Program (JEP), has three objectives: (1) provide constructive feedback to judges, executive committees and national federations; (2) assign the best judges to the most important competitions; (3) detect bias and outright cheating. Using data from international gymnastics competitions held during the 2013-2016 Olympic cycle, we first develop a marking score evaluating the accuracy of the marks given by gymnastics judges. Judging a gymnastics routine is a random process, and we can model this process very accurately using heteroscedastic random variables. The marking score scales the difference between the mark of a judge and the theoretical performance of a gymnast as a function of the standard deviation of the judging error estimated from data for each apparatus. This dependence between judging variability and performance quality has never been properly studied. We then study ranking scores assessing to what extent judges rate gymnasts in the correct order, and explain why we ultimately chose not to implement them. We also study outlier detection to pinpoint gymnasts who were poorly evaluated by judges. Finally, we discuss interesting observations and discoveries that led to recommendations and rule changes at the FIG.

研究动机与目标

  • 開發一種客觀、數據驅動的方法,用於評估大型賽事中國際體操裁判的準確性與一致性。
  • 透過分析裁判評分與預期表現水平之間的偏差,檢測系統性偏見或潛在作弊行為。
  • 為裁判、國家體協與執行委員會提供具體反饋,以改善培訓與資格認證。
  • 支援在奧運等高風險賽事中選拔表現優異的裁判。
  • 透過標準化指引提升賽後評估中控制分數的可靠性。

提出的方法

  • 評分機制將評分過程建模為異方差隨機過程,其中變異性取決於各器械的內在誤差,從而實現針對各器械的準確性評估。
  • 控制分數透過現場賽事中所有裁判評分的中位數近似估算,賽後則使用更精確的影片審核分數進行驗證。
  • 長期追蹤分析涵蓋數千次評分,用以評估裁判在時間軸上的穩定性與準確性。
  • 異常值檢測可識別出在評分面板中獲得顯著不一致或不準確分數的體操選手。
  • 排名分數雖經評估,但因不穩定且對微小分數差異過度敏感,最終未予實施。
  • 統計模型使用2013至2016年奧運周期的數據進行訓練,當官方控制分數不可得時,以中位數分數作為替代指標。

实验结果

研究问题

  • RQ1如何利用統計模型客觀衡量國際體操裁判的準確性?
  • RQ2各器械之間內在評分誤差變異程度差異有多大?其對評分一致性有何影響?
  • RQ3統計引擎能否在實時與回溯評分中檢測偏見或作弊行為?
  • RQ4現場賽事中以中位數分數作為控制分數近似值,如何影響裁判評估的可靠性?
  • RQ5長期表現追蹤在識別持續準確或不一致的裁判方面扮演何種角色?

主要发现

  • 評分機制透過納入器械特異的內在評分誤差變異性,有效量化裁判的準確性,進而實現裁判間更公平的比較。
  • 即使在Category 1裁判中,仍觀察到顯著的表現差異,顯示並非所有頂尖裁判都具有同等準確性。
  • 異常值檢測成功識別出評分不當的體操選手,凸顯評分小組評分不一致的潛在問題。
  • 現場賽事中以中位數分數作為控制分數的代理指標,雖提供合理近似,但若評分小組系統性不準確,可能造成誤導。
  • 賽後由技術委員會進行的影片審核仍為黃金標準,但其可靠性取決於委員會組成的公正性與一致的整合方法。
  • 本研究促成國際體操聯合會(FIG)提出建議與規則修訂,包括改進控制分數計算指引,並加強對評分表現的監控。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。