[论文解读] PEg TRAnsfer Workflow recognition challenge report: Does multi-modal data improve recognition?
本论文评估了多模态数据(特别是视频和运动学数据)在受控 peg 转移任务中对外科手术流程识别的影响。基于包含150个序列的基准数据集,其标注粒度分为三个层次,研究结果表明,将视频与运动学数据结合可显著提升识别准确率(AD-Accuracy 最高达93%),相较于单模态方法,但计算成本大幅增加。
This paper presents the design and results of the "PEg TRAnsfert Workflow recognition" (PETRAW) challenge whose objective was to develop surgical workflow recognition methods based on one or several modalities, among video, kinematic, and segmentation data, in order to study their added value. The PETRAW challenge provided a data set of 150 peg transfer sequences performed on a virtual simulator. This data set was composed of videos, kinematics, semantic segmentation, and workflow annotations which described the sequences at three different granularity levels: phase, step, and activity. Five tasks were proposed to the participants: three of them were related to the recognition of all granularities with one of the available modalities, while the others addressed the recognition with a combination of modalities. Average application-dependent balanced accuracy (AD-Accuracy) was used as evaluation metric to take unbalanced classes into account and because it is more clinically relevant than a frame-by-frame score. Seven teams participated in at least one task and four of them in all tasks. Best results are obtained with the use of the video and the kinematics data with an AD-Accuracy between 93% and 90% for the four teams who participated in all tasks. The improvement between video/kinematic-based methods and the uni-modality ones was significant for all of the teams. However, the difference in testing execution time between the video/kinematic-based and the kinematic-based methods has to be taken into consideration. Is it relevant to spend 20 to 200 times more computing time for less than 3% of improvement? The PETRAW data set is publicly available at www.synapse.org/PETRAW to encourage further research in surgical workflow recognition.
研究动机与目标
- 探究结合多种数据模态是否能提升外科手术流程识别的准确率,相较于单模态方法。
- 开发一个标准化基准数据集,用于外科手术流程识别,涵盖视频、运动学和语义分割等多种模态。
- 在三个粒度层次上评估识别性能:阶段、步骤和活动。
- 评估使用多模态输入时,识别准确率提升与计算成本增加之间的权衡。
- 通过公开发布 PETRAW 数据集,推动开放科学,支持未来在手术过程建模领域的研究。
提出的方法
- 本研究组织了 PEg TRAnsfer Workflow(PETRAW)挑战赛,提供了来自虚拟外科模拟器的150个 peg 转移序列的数据集。
- 该数据集包含同步的视频、运动学轨迹、语义分割掩码以及多级标注(阶段、步骤、活动)。
- 定义了五个任务:三个单模态任务(仅视频、仅运动学、仅语义分割)和两个多模态任务(视频+运动学、视频+分割)。
- 参赛者应用深度学习模型——主要为卷积神经网络(CNN)、循环神经网络(RNN)和 Transformer——在各自模态上进行训练,以识别流程阶段。
- 评估采用应用相关平衡准确率(AD-Accuracy),该指标考虑了类别不平衡问题,相较于帧级 F1 分数更具临床相关性。
- 结果在所有序列上聚合,并按任务进行评估,以比较单模态与多模态方法的性能及计算效率。
实验结果
研究问题
- RQ1与单独使用任一模态相比,结合视频和运动学数据是否能显著提升外科手术流程识别的准确率?
- RQ2语义分割数据的引入如何影响不同粒度层次下的识别性能?
- RQ3在使用多模态输入时,识别准确率与计算成本之间的权衡如何?
- RQ4多模态模型是否能在不同外科任务和不同程序抽象层次上实现良好泛化?
- RQ5在单模态与多模态设置下,不同深度学习架构在手术流程识别中的表现如何?
主要发现
- 表现最佳的模型结合了视频与运动学数据,在参与全部任务的四支团队中,实现了93%至90%的应用相关平衡准确率(AD-Accuracy)。
- 多模态方法始终优于单模态方法,其中视频+运动学组合相较于仅使用视频或仅使用运动学的基线模型表现提升最为显著。
- 尽管在某些情况下准确率仅提升3%,使用视频与运动学数据的计算成本相比仅使用运动学的方法增加了2000%至20000%。
- 仅使用语义分割数据未能取得具有竞争力的结果,表明在该设置下其在流程识别中独立应用的潜力有限。
- PETRAW 数据集已公开发布于 www.synapse.org/PETRAW,可支持未来在手术流程识别研究中的基准测试与可复现性研究。
- 本研究证实,多模态融合可显著提升识别的鲁棒性与准确率,尤其在步骤和活动等更细粒度层次上表现突出,但计算效率仍是关键制约因素。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。