[Paper Review] PEg TRAnsfer Workflow recognition challenge report: Does multi-modal data improve recognition?
This paper evaluates the impact of multi-modal data—specifically video and kinematic data—on surgical workflow recognition in a controlled peg transfer task. Using a benchmark dataset of 150 sequences with annotations at three granularities, the study demonstrates that combining video and kinematics significantly improves recognition accuracy (AD-Accuracy up to 93%) compared to unimodal approaches, though with substantial increases in computational cost.
This paper presents the design and results of the "PEg TRAnsfert Workflow recognition" (PETRAW) challenge whose objective was to develop surgical workflow recognition methods based on one or several modalities, among video, kinematic, and segmentation data, in order to study their added value. The PETRAW challenge provided a data set of 150 peg transfer sequences performed on a virtual simulator. This data set was composed of videos, kinematics, semantic segmentation, and workflow annotations which described the sequences at three different granularity levels: phase, step, and activity. Five tasks were proposed to the participants: three of them were related to the recognition of all granularities with one of the available modalities, while the others addressed the recognition with a combination of modalities. Average application-dependent balanced accuracy (AD-Accuracy) was used as evaluation metric to take unbalanced classes into account and because it is more clinically relevant than a frame-by-frame score. Seven teams participated in at least one task and four of them in all tasks. Best results are obtained with the use of the video and the kinematics data with an AD-Accuracy between 93% and 90% for the four teams who participated in all tasks. The improvement between video/kinematic-based methods and the uni-modality ones was significant for all of the teams. However, the difference in testing execution time between the video/kinematic-based and the kinematic-based methods has to be taken into consideration. Is it relevant to spend 20 to 200 times more computing time for less than 3% of improvement? The PETRAW data set is publicly available at www.synapse.org/PETRAW to encourage further research in surgical workflow recognition.
Motivation & Objective
- To investigate whether combining multiple data modalities improves surgical workflow recognition accuracy compared to unimodal approaches.
- To develop a standardized benchmark dataset for surgical workflow recognition using diverse modalities including video, kinematics, and semantic segmentation.
- To evaluate recognition performance across three levels of granularity: phase, step, and activity.
- To assess the trade-off between improved accuracy and increased computational cost when using multi-modal inputs.
- To promote open science by releasing the PETRAW dataset for future research in surgical process modeling.
Proposed method
- The study organized the PEg TRAnsfer Workflow (PETRAW) challenge, providing a dataset of 150 peg transfer sequences from a virtual surgical simulator.
- The dataset includes synchronized video, kinematic trajectories, semantic segmentation masks, and multi-level annotations (phase, step, activity).
- Five tasks were defined: three unimodal (video-only, kinematics-only, semantic segmentation-only) and two multimodal (video + kinematics, video + segmentation).
- Participants applied deep learning models—primarily CNNs, RNNs, and Transformers—trained on the respective modalities to recognize workflow stages.
- Evaluation used application-dependent balanced accuracy (AD-Accuracy), which accounts for class imbalance and is more clinically relevant than frame-level F1 scores.
- Results were aggregated across all sequences and evaluated per task to compare unimodal vs. multimodal performance and computational efficiency.
Experimental results
Research questions
- RQ1Does combining video and kinematic data significantly improve surgical workflow recognition accuracy compared to using either modality alone?
- RQ2How does the inclusion of semantic segmentation data affect recognition performance across different levels of granularity?
- RQ3What is the trade-off between recognition accuracy and computational cost when using multi-modal inputs?
- RQ4Can multi-modal models generalize across different surgical tasks and levels of procedural abstraction?
- RQ5To what extent do different deep learning architectures perform across unimodal and multimodal settings in surgical workflow recognition?
Key findings
- The best-performing models combined video and kinematic data, achieving an application-dependent balanced accuracy (AD-Accuracy) between 93% and 90% across the four teams that participated in all tasks.
- Multi-modal methods consistently outperformed unimodal approaches, with video+kinematics showing the most significant improvement over video-only or kinematics-only baselines.
- The use of video and kinematics increased computational cost by 2,000% to 20,000% compared to kinematics-only methods, despite only a 3% accuracy gain in some cases.
- Semantic segmentation alone did not yield competitive results compared to video or kinematics, suggesting limited standalone utility for workflow recognition in this setting.
- The PETRAW dataset is publicly available at www.synapse.org/PETRAW, enabling future benchmarking and reproducibility in surgical workflow recognition research.
- The study confirms that multi-modal fusion enhances recognition robustness and accuracy, particularly at finer granularities like step and activity, though efficiency remains a critical constraint.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.