[Paper Review] Next-generation Surgical Navigation: Marker-less Multi-view 6DoF Pose Estimation of Surgical Instruments
This paper presents a marker-less, multi-view 6DoF pose estimation system for surgical instruments using RGB-D cameras, achieving sub-millimeter accuracy (1.01 mm position, 0.89° orientation) with five cameras under optimal conditions. It introduces a novel multi-view RGB-D dataset from ex-vivo spine surgeries and demonstrates that multi-view fusion significantly improves accuracy and occlusion robustness over single-view methods, making marker-less tracking a viable alternative to traditional systems.
State-of-the-art research of traditional computer vision is increasingly leveraged in the surgical domain. A particular focus in computer-assisted surgery is to replace marker-based tracking systems for instrument localization with pure image-based 6DoF pose estimation using deep-learning methods. However, state-of-the-art single-view pose estimation methods do not yet meet the accuracy required for surgical navigation. In this context, we investigate the benefits of multi-view setups for highly accurate and occlusion-robust 6DoF pose estimation of surgical instruments and derive recommendations for an ideal camera system that addresses the challenges in the operating room. The contributions of this work are threefold. First, we present a multi-camera capture setup consisting of static and head-mounted cameras, which allows us to study the performance of pose estimation methods under various camera configurations. Second, we publish a multi-view RGB-D video dataset of ex-vivo spine surgeries, captured in a surgical wet lab and a real operating theatre and including rich annotations for surgeon, instrument, and patient anatomy. Third, we evaluate three state-of-the-art single-view and multi-view methods for the task of 6DoF pose estimation of surgical instruments and analyze the influence of camera configurations, training data, and occlusions on the pose accuracy and generalization ability. The best method utilizes five cameras in a multi-view pose optimization and achieves an average position and orientation error of 1.01 mm and 0.89° for a surgical drill as well as 2.79 mm and 3.33° for a screwdriver under optimal conditions. Our results demonstrate that marker-less tracking of surgical instruments is becoming a feasible alternative to existing marker-based systems.
Motivation & Objective
- To develop a marker-less, high-accuracy 6DoF pose estimation system for surgical instruments using multi-view computer vision.
- To address limitations of marker-based systems, such as line-of-sight constraints, calibration complexity, and workflow disruption.
- To evaluate the impact of camera configuration, training data, and occlusions on pose accuracy and generalization in surgical environments.
- To provide a publicly available multi-view RGB-D dataset for training and benchmarking in surgical navigation.
- To demonstrate that multi-view fusion enables millimeter-level accuracy and robustness under real-world surgical conditions.
Proposed method
- Deployed a multi-camera setup combining static and head-mounted cameras (e.g., Azure Kinect, HoloLens 2) to capture synchronized RGB-D video from multiple perspectives.
- Developed a multi-view pose optimization framework that fuses 2D-3D correspondences across views to improve 6DoF pose estimation accuracy.
- Utilized both real ex-vivo surgical data and synthetic data for training, with data augmentation to improve robustness to occlusions and lighting variations.
- Applied state-of-the-art 6DoF pose estimation networks (e.g., from BOP benchmark) and adapted them for surgical instrument tracking in multi-view settings.
- Performed in-domain fine-tuning using real test-time data to further reduce pose errors and improve generalization.
- Annotated the dataset with ground truth poses for instruments, surgeons, and patient anatomy, enabling precise evaluation.
Experimental results
Research questions
- RQ1Can multi-view RGB-D fusion achieve sub-millimeter accuracy in 6DoF pose estimation of surgical instruments under realistic operating room conditions?
- RQ2How do different camera configurations (number, placement, type) affect pose accuracy and robustness to occlusions?
- RQ3To what extent does synthetic data improve model generalization when training on real surgical data?
- RQ4Can marker-less tracking systems outperform or match the accuracy of existing marker-based navigation systems in clinical-relevant scenarios?
- RQ5What camera setup and data strategy are optimal for achieving high accuracy and robustness in real-time surgical navigation?
Key findings
- The best-performing multi-view method using five cameras achieved a mean position error of 1.01 mm and orientation error of 0.89° for a surgical drill under optimal conditions.
- For a screwdriver, the method achieved 2.79 mm position error and 3.33° orientation error, demonstrating robustness across different instrument types.
- Even with only two cameras, the system achieved millimeter-level accuracy, proving that marker-less tracking is feasible as a viable alternative to marker-based systems.
- In-domain fine-tuning on real test-time data reduced pose errors significantly, highlighting the value of domain-specific adaptation.
- Synthetic data was shown to improve model robustness, especially in handling complex occlusions and lighting variations common in the OR.
- The proposed multi-view setup maintained functionality even when marker-based systems failed due to instrument out-of-view, demonstrating superior occlusion resilience.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.