Skip to main content
QUICK REVIEW

[论文解读] Prediction of the FIFA World Cup 2018 - A random forest approach with an emphasis on estimated team ability parameters

Andreas Groll, Christophe Ley|arXiv (Cornell University)|Jun 8, 2018
Sports Analytics and Performance参考文献 20被引用 9
一句话总结

本文提出了一种结合团队实力参数的混合随机森林模型,通过排名方法估算团队实力参数,以预测2018年国际足联世界杯赛果。通过融合机器学习与基于实力的团队实力估计方法,该模型在预测性能上优于传统的泊松回归和独立的随机森林模型,预测西班牙为最有可能的冠军,略微领先于卫冕冠军德国。

ABSTRACT

In this work, we compare three different modeling approaches for the scores of soccer matches with regard to their predictive performances based on all matches from the four previous FIFA World Cups 2002 - 2014: Poisson regression models, random forests and ranking methods. While the former two are based on the teams' covariate information, the latter method estimates adequate ability parameters that reflect the current strength of the teams best. Within this comparison the best-performing prediction methods on the training data turn out to be the ranking methods and the random forests. However, we show that by combining the random forest with the team ability parameters from the ranking methods as an additional covariate we can improve the predictive power substantially. Finally, this combination of methods is chosen as the final model and based on its estimates, the FIFA World Cup 2018 is simulated repeatedly and winning probabilities are obtained for all teams. The model slightly favors Spain before the defending champion Germany. Additionally, we provide survival probabilities for all teams and at all tournament stages as well as the most probable tournament outcome.

研究动机与目标

  • 使用2002–2014年国际足联世界杯数据,比较泊松回归、随机森林和排名方法在足球比赛得分预测中的性能。
  • 评估将排名方法估算的团队实力参数整合到随机森林模型中,是否能显著提升预测准确性。
  • 使用表现最佳的模型对2018年国际足联世界杯进行10万次模拟,以估算夺冠概率、生存率及最可能的赛事路径。
  • 提供超越博彩公司赔率的稳健、数据驱动的球队表现与赛事进程预测。

提出的方法

  • 本研究使用基于2002–2014年国际足联世界杯比赛级协变量训练的随机森林模型,包括球队特定特征,如国际足联排名、晋级表现以及历史世界杯成绩。
  • 团队实力参数通过一种排名方法估算,该方法从历史比赛结果中推断当前球队实力,作为额外的协变量输入。
  • 最终模型将这些估算的团队实力参数作为补充特征整合进随机森林算法,以增强预测能力。
  • 模型使用斯凯勒姆分布对2018年赛事进行10万次模拟,以建模比赛结果,包括淘汰赛阶段平局时的加时赛和点球大战。
  • 在每个赛事阶段计算所有球队的生存概率,并从模拟集合中推导出最可能的赛事进程路径。
  • 使用先前赛事的训练数据评估预测性能,模型比较基于预测准确率指标。

实验结果

研究问题

  • RQ1在三种建模方法中——泊松回归、随机森林和排名方法——哪一种对国际足联世界杯足球比赛得分的预测准确率最高?
  • RQ2与独立模型相比,将排名方法估算的团队实力参数整合进随机森林模型,是否能显著提升预测性能?
  • RQ3基于表现最佳的模型,2018年国际足联世界杯各队的夺冠概率是多少?
  • RQ4生存概率在不同赛事阶段如何变化?每支队伍最可能的晋级路径是什么?
  • RQ5为何该模型预测西班牙为最有可能的冠军,尽管德国是博彩公司和专家的首选?

主要发现

  • 尽管德国是博彩公司的首选,但该模型预测西班牙为2018年国际足联世界杯最有可能的冠军,其夺冠概率高于德国。
  • 结合排名方法估算的团队实力参数的混合随机森林模型,在预测性能上优于独立的泊松回归、随机森林和排名方法。
  • 德国有61%的概率晋级十六强,但其整体夺冠概率低于西班牙,主要因其早期被淘汰的风险更高。
  • 在成功晋级八分之一决赛的前提下,德国成为对西班牙的热门,表明德国的主要弱点在于淘汰赛早期阶段。
  • 最可能的赛事路径(即德国夺冠)的总体概率极低,仅为1.55 × 10−5%,凸显了赛事结果的高度不确定性。
  • 为所有球队在每个阶段提供了生存概率,同时基于模拟提供了晋级可能性和比赛结果的详细估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。