Skip to main content
QUICK REVIEW

[论文解读] Using Twitter to predict football outcomes

Stylianos Kampakis, Andreas Adamides|arXiv (Cornell University)|Nov 5, 2014
Sports Analytics and Performance参考文献 8被引用 12
一句话总结

本研究探讨了Twitter数据是否能够预测英超联赛足球比赛结果。通过使用自然语言处理(NLP)和情感分析对推文进行分析,作者构建了预测模型,其表现优于随机猜测,且与仅使用历史统计数据的模型性能相当,而组合模型则展现出更高的准确性。

ABSTRACT

Twitter has been proven to be a notable source for predictive modelling on various domains such as the stock market, the dissemination of diseases or sports outcomes. However, such a study has not been conducted in football (soccer) so far. The purpose of this research was to study whether data mined from Twitter can be used for this purpose. We built a set of predictive models for the outcome of football games of the English Premier League for a 3 month period based on tweets and we studied whether these models can overcome predictive models which use only historical data and simple football statistics. Moreover, combined models are constructed using both Twitter and historical data. The final results indicate that data mined from Twitter can indeed be a useful source for predicting games in the Premier League. The final Twitter-based model performs significantly better than chance when measured by Cohen's kappa and is comparable to the model that uses simple statistics and historical data. Combining both models raises the performance higher than it was achieved by each individual model. Thereby, this study provides evidence that Twitter derived features can indeed provide useful information for the prediction of football (soccer) outcomes.

研究动机与目标

  • 确定Twitter数据是否可作为英格兰英超联赛足球比赛结果的可靠预测指标。
  • 比较基于Twitter数据的特征与仅依赖历史足球统计数据的模型之间的预测能力。
  • 评估将Twitter数据与传统统计数据结合是否能提升预测准确性。
  • 评估实时社交媒体情感在体育结果预测中的可行性。

提出的方法

  • 使用Twitter API在为期3个月的时间内收集与英超联赛比赛相关的推文。
  • 应用自然语言处理(NLP)技术,从推文中提取情感和主题特征。
  • 基于Twitter提取的特征,使用逻辑回归及其他机器学习分类器构建预测模型。
  • 仅使用历史比赛数据和简单足球统计数据(如球队排名、进球差)构建基线模型。
  • 通过将Twitter特征与历史统计数据结合,构建混合模型以提升预测性能。
  • 使用Cohen's kappa和准确率指标评估模型性能,并与随机猜测和基线模型进行比较。

实验结果

研究问题

  • RQ1Twitter情感和讨论量能否预测英格兰英超联赛足球比赛的结果?
  • RQ2基于Twitter的模型预测表现与仅使用历史足球统计数据的模型相比如何?
  • RQ3将Twitter数据与传统统计数据结合,能在多大程度上提升预测准确性?
  • RQ4Twitter数据的预测能力是否在统计上显著高于随机猜测?

主要发现

  • 基于Twitter的预测模型在Cohen's kappa指标下显著优于随机猜测。
  • 仅使用Twitter数据的模型表现与仅使用历史足球统计数据的模型相当。
  • 结合Twitter数据与历史统计数据的混合模型,其预测准确性高于单独使用任一数据源的模型。
  • Twitter上的情感倾向和讨论量被证实是比赛结果的有效预测指标。
  • 本研究证明,实时社交媒体数据可在体育预测中提供有价值的预测洞察。
  • 研究结果支持将社交媒体作为体育分析中互补数据源的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。