[论文解读] Vid2Player: Controllable Video Sprites that Behave and Appear like Professional Tennis Players
Vid2Player 通过从标注的转播画面中建模球员的行为与外观,生成交互式、逼真的职业网球选手视频精灵。它利用击球周期状态机和数据驱动的行为模型来预测击球落点与场上站位,从而合成新颖且合理的网球回合——包括虚构的对决——实现在大规模场景下的交互控制与逼真对打动态。
We present a system that converts annotated broadcast video of tennis matches into interactively controllable video sprites that behave and appear like professional tennis players. Our approach is based on controllable video textures, and utilizes domain knowledge of the cyclic structure of tennis rallies to place clip transitions and accept control inputs at key decision-making moments of point play. Most importantly, we use points from the video collection to model a player's court positioning and shot selection decisions during points. We use these behavioral models to select video clips that reflect actions the real-life player is likely to take in a given match play situation, yielding sprites that behave realistically at the macro level of full points, not just individual tennis motions. Our system can generate novel points between professional tennis players that resemble Wimbledon broadcasts, enabling new experiences such as the creation of matchups between players that have not competed in real life, or interactive control of players in the Wimbledon final. According to expert tennis players, the rallies generated using our approach are significantly more realistic in terms of player behavior than video sprite methods that only consider the quality of motion transitions during video synthesis.
研究动机与目标
- 生成在行为上如同真实运动员的交互式、视觉逼真的职业网球选手视频精灵。
- 利用真实比赛数据建模球员行为(如击球选择与场上站位),以确保战术上的真实性。
- 实现从未实际交手过的球员之间的新型网球回合,包括虚构对决。
- 在回合合成过程中支持用户对击球落点与恢复站位的交互式控制。
- 通过整合宏观层面的战术行为,将视频精灵的真实感提升至超越运动质量的层面。
提出的方法
- 该系统使用击球周期状态机,在单个击球层级上组织视频合成,确保在回合关键决策点发生过渡。
- 从标注的广播视频中构建球员特定的行为模型,基于对手和比赛情境预测击球落点与场上站位。
- 通过可控的视频纹理系统,根据视觉质量和预测行为选择视频片段,确保动作与策略的逼真性。
- 利用神经图像到图像的翻译技术,校正不同比赛日之间的外观差异,并处理单视角广播画面中的部分遮挡。
- 通过领域感知的视频合成,增强对噪声计算机视觉标注(如姿态估计误差)的鲁棒性。
- 交互控制通过用户输入实现,用户可设定期望的击球落点(红色圆点)与恢复位置(蓝色方块),实时引导行为模型。
实验结果
研究问题
- RQ1是否能够生成不仅模仿真实运动,还能表现出真实职业网球选手战术行为的视频精灵?
- RQ2如何利用网球回合结构的领域知识,以提升视频精灵合成的真实感与效率?
- RQ3数据驱动的击球选择与场上站位行为模型在多大程度上能增强合成网球回合的宏观层面真实感?
- RQ4该系统能否生成在现实中从未交手过的球员之间合理且新颖的网球比赛?
- RQ5交互式用户控制在引导视频精灵行为时,其在视觉与战术合理性上的有效性如何?
主要发现
- 该系统成功生成了真实球员之间此前未匹配过的新型网球回合,例如罗杰·费德勒对塞雷娜·威廉姆斯,其视觉与行为表现均显得自然可信。
- 专业网球选手评价 Vid2Player 生成的回合在球员行为方面显著比仅关注运动过渡质量的基线视频精灵方法更真实。
- 行为模型揭示了显著的战术差异:诺瓦克·德约维奇比对纳达尔更频繁地针对费德勒的弱侧反手进行攻击,这与真实比赛中的战术倾向一致。
- 系统支持交互式控制,用户可通过点击设定击球落点与恢复位置目标,球员视频精灵能实时据此调整行为。
- 使用击球周期状态机显著降低了片段搜索成本,并通过仅在低运动量、高稳定性时刻进行过渡,提升了视觉质量。
- 通过神经图像翻译与幻觉生成技术,系统在面对标注误差与外观差异时表现出鲁棒性,保持了在多样化广播片段间的一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。