[Paper Review] What is the most competitive sport?
This paper proposes the upset frequency $ q $, the likelihood of underdog victories, as a quantitative measure of competitiveness in sports. Using a stochastic model that links team parity (variance in winning fractions) to predictability, the authors analyze 100+ years of data across five major leagues and find that soccer and baseball are the most competitive, while football and basketball are the least, with $ q = 0.452 $ and $ 0.441 $, respectively.
We present an extensive statistical analysis of the results of all sports competitions in five major sports leagues in England and the United States. We characterize the parity among teams by the variance in the winning fraction from season-end standings data and quantify the predictability of games by the frequency of upsets from game results data. We introduce a mathematical model in which the underdog team wins with a fixed upset probability. This model quantitatively relates the parity among teams with the predictability of the games, and it can be used to estimate the upset frequency from standings data. We propose the likelihood of upsets as a measure of competitiveness.
Motivation & Objective
- To develop a quantitative measure of competitiveness in professional sports that accounts for varying season lengths and team parity.
- To investigate the relationship between team performance variance (parity) and game predictability (upset frequency).
- To test whether the frequency of upsets ($ q $) can serve as a robust, standalone index of competitiveness across different sports.
- To validate a theoretical model linking standings data (winning fraction variance) to game-level outcomes (upset frequency).
Proposed method
- The authors define the upset frequency $ q $ as the fraction of games won by the team with the worse record at game time, excluding ties and early-season games.
- They introduce a stochastic model where the underdog wins with fixed probability $ q < 0.5 $, and the favorite wins otherwise, simulating game outcomes based on relative team strength.
- Monte Carlo simulations of artificial leagues are used to fit the model to real data by matching the simulated distribution of winning fractions to actual season-end standings.
- The model predicts that in the limit of infinite games, the standard deviation $ ho $ of winning fractions is $ (1/2 - q)/ ho{3} $, linking $ ho $ and $ q $ analytically.
- Theoretical predictions for $ q_{\text{model}} $ are derived from fitting the simulated distribution to observed $ F(x) $, the cumulative distribution of winning fractions.
- The analysis accounts for tie handling by assigning 0.5 points to each team in tied games, and verifies that this has minimal impact on $ q $.
Experimental results
Research questions
- RQ1How does the frequency of upsets ($ q $) vary across major professional sports leagues, and can it serve as a reliable measure of competitiveness?
- RQ2To what extent does the length of a season influence the observed variance in team winning fractions, and how can this be corrected for fair comparison?
- RQ3Can a simple stochastic model with a fixed upset probability $ q $ accurately reproduce the observed distribution of team winning fractions across different sports?
- RQ4Is there a consistent relationship between team parity (measured by $ ho $) and predictability (measured by $ q $), and can this be used to estimate $ q $ from standings data alone?
- RQ5How have trends in $ q $ and $ ho $ evolved over time in different leagues, and what might they indicate about changes in team balance or game strategy?
Key findings
- Soccer (FA) and baseball (MLB) exhibit the highest upset frequencies at $ q = 0.452 $ and $ q = 0.441 $, respectively, indicating they are the most competitive sports.
- Basketball (NBA) and football (NFL) have the lowest upset frequencies at $ q = 0.365 $ and $ q = 0.364 $, respectively, indicating they are the least competitive.
- The model-predicted upset probability $ q_{\text{model}} $ closely matches the measured $ q $, with differences of less than 0.05 across all leagues, validating the model’s accuracy.
- The variance in winning fractions $ ho $ is inversely related to the bias $ 1/2 - q $, confirming the theoretical prediction that higher $ q $ leads to lower $ ho $.
- Over the past 60 years, NFL and MLB games have become more competitive, as shown by increasing $ q $, while the English Football Association (FA) has shown a declining trend in $ q $.
- The model enables estimation of $ q $ from standings data alone, providing a practical tool for assessing competitiveness without access to game-by-game results.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.