[论文解读] ShuttleSet: A Human-Annotated Stroke-Level Singles Dataset for Badminton Tactical Analysis
ShuttleSet 是一个大规模、人工标注的羽毛球单打逐拍数据集,包含 44 场比赛、3,685 个回合和 36,492 次击球,数据来自 2018–2021 年间排名靠前的选手。标注人员使用计算机辅助标注工具记录了 18 种击球类型、击球位置及击球时的球员位置,支持高级战术分析,并作为先进模型在击球预测、影响力分析和移动预测方面的基准。
With the recent progress in sports analytics, deep learning approaches have demonstrated the effectiveness of mining insights into players' tactics for improving performance quality and fan engagement. This is attributed to the availability of public ground-truth datasets. While there are a few available datasets for turn-based sports for action detection, these datasets severely lack structured source data and stroke-level records since these require high-cost labeling efforts from domain experts and are hard to detect using automatic techniques. Consequently, the development of artificial intelligence approaches is significantly hindered when existing models are applied to more challenging structured turn-based sequences. In this paper, we present ShuttleSet, the largest publicly-available badminton singles dataset with annotated stroke-level records. It contains 104 sets, 3,685 rallies, and 36,492 strokes in 44 matches between 2018 and 2021 with 27 top-ranking men's singles and women's singles players. ShuttleSet is manually annotated with a computer-aided labeling tool to increase the labeling efficiency and effectiveness of selecting the shot type with a choice of 18 distinct classes, the corresponding hitting locations, and the locations of both players at each stroke. In the experiments, we provide multiple benchmarks (i.e., stroke influence, stroke forecasting, and movement forecasting) with baselines to illustrate the practicability of using ShuttleSet for turn-based analytics, which is expected to stimulate both academic and sports communities. Over the past two years, a visualization platform has been deployed to illustrate the variability of analysis cases from ShuttleSet for coaches to delve into players' tactical preferences with human-interactive interfaces, which was also used by national badminton teams during multiple international high-ranking matches.
研究动机与目标
- 为像羽毛球这样的回合制运动解决缺乏公开可用、高质量、逐拍标注数据集的问题。
- 为精英级别单打比赛中的每一次击球提供结构化、微观级别的元数据(击球类型、位置、球员位置)。
- 支持体育分析中高级机器学习与深度学习应用,特别是战术模式识别与预测。
- 通过提供一个交互式可视化平台,弥合学术研究与体育实践之间的差距,使教练和分析师能够无需技术专长即可探索球员战术。
- 利用真实世界精英级别羽毛球数据,建立击球影响力、击球预测和移动预测的基准。
提出的方法
- 通过领域专家使用 S2-标注工具进行人工标注构建数据集,该工具支持对逐拍事件进行高效且一致的标注。
- 每记击球均使用 BLSR 格式进行结构化数据表示,标注了 18 种不同的击球类型、击球位置以及击球瞬间的球员位置。
- 计算机辅助标注界面减少了标注时间并提高了标注一致性,降低了在复杂高速序列中的误标风险。
- 数据集涵盖 2018–2021 年间 44 场比赛,涉及 27 名男女单打世界排名前列的选手,共 104 局和 3,685 个回合。
- 建立了多个基准:击球影响力(分类任务)、击球预测(序列建模)和移动预测(时空预测)。
- 开发了一个交互式可视化平台,使教练和分析师能够利用数据集中的真实比赛数据,无需技术背景即可探索战术模式。

实验结果
研究问题
- RQ1羽毛球逐拍标注在提升机器学习模型在战术预测与动作识别方面的性能方面有何作用?
- RQ2精英男、女单打选手在站位和击球选择模式上存在哪些关键差异?
- RQ3基于历史击球与位置数据,基于序列的模型在多大程度上能够预测下一次击球或移动?
- RQ4战术模式(如网前搓球使用与场地覆盖)在不同比赛阶段(如半决赛与决赛)有何差异?
- RQ5基于逐拍数据集构建的交互式可视化平台,能否有效支持教练分析球员战术与对手倾向?
主要发现
- ShuttleSet 是目前公开可用的、拥有逐拍标注的最大羽毛球单打数据集,涵盖 44 场精英级别比赛中的 36,492 次击球。
- 数据集包含 18 种不同的击球类型、精确的击球位置以及每次击球时的球员位置,支持精细化战术分析。
- 基准测试结果表明,当前模型在击球影响力、预测与移动预测任务上表现尚可但未达最优,表明仍有改进空间。
- 对维克多·阿克塞尔森与桃田贤斗比赛的分析显示,保持中心位置的稳定性以及有效使用网前搓球,与前场对抗中的成功率呈正相关。
- 可视化平台成功帮助国家队在国际比赛中分析对手战术,展现出实际应用价值。
- 该数据集与基准有望推动回合制体育分析中序列建模、图神经网络方法及球员风格对比研究的发展。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。