Skip to main content
QUICK REVIEW

[论文解读] Next-generation Surgical Navigation: Marker-less Multi-view 6DoF Pose Estimation of Surgical Instruments

Jonas Hein, Nicola Cavalcanti|arXiv (Cornell University)|May 5, 2023
Surgical Simulation and Training被引用 4
一句话总结

本文提出了一种基于RGB-D相机的无标记多视角6自由度(6DoF)手术器械位姿估计系统,在理想条件下使用五台相机实现了亚毫米级精度(位置误差1.01 mm,姿态误差0.89°)。该研究构建了一个来自离体脊柱手术的新型多视角RGB-D数据集,并证明多视角融合显著提升了精度和遮挡鲁棒性,使无标记跟踪成为传统系统的一种可行替代方案。

ABSTRACT

State-of-the-art research of traditional computer vision is increasingly leveraged in the surgical domain. A particular focus in computer-assisted surgery is to replace marker-based tracking systems for instrument localization with pure image-based 6DoF pose estimation using deep-learning methods. However, state-of-the-art single-view pose estimation methods do not yet meet the accuracy required for surgical navigation. In this context, we investigate the benefits of multi-view setups for highly accurate and occlusion-robust 6DoF pose estimation of surgical instruments and derive recommendations for an ideal camera system that addresses the challenges in the operating room. The contributions of this work are threefold. First, we present a multi-camera capture setup consisting of static and head-mounted cameras, which allows us to study the performance of pose estimation methods under various camera configurations. Second, we publish a multi-view RGB-D video dataset of ex-vivo spine surgeries, captured in a surgical wet lab and a real operating theatre and including rich annotations for surgeon, instrument, and patient anatomy. Third, we evaluate three state-of-the-art single-view and multi-view methods for the task of 6DoF pose estimation of surgical instruments and analyze the influence of camera configurations, training data, and occlusions on the pose accuracy and generalization ability. The best method utilizes five cameras in a multi-view pose optimization and achieves an average position and orientation error of 1.01 mm and 0.89° for a surgical drill as well as 2.79 mm and 3.33° for a screwdriver under optimal conditions. Our results demonstrate that marker-less tracking of surgical instruments is becoming a feasible alternative to existing marker-based systems.

研究动机与目标

  • 开发一种基于多视角计算机视觉的无标记、高精度6DoF手术器械位姿估计系统。
  • 解决标记系统存在的局限性,如视线约束、标定复杂性和工作流程干扰。
  • 评估相机配置、训练数据和遮挡对手术环境中位姿精度和泛化能力的影响。
  • 提供一个公开可用的多视角RGB-D数据集,用于手术导航中的训练和基准测试。
  • 证明多视角融合可在真实手术条件下实现毫米级精度和鲁棒性。

提出的方法

  • 部署了结合静态相机和头戴式相机(如Azure Kinect、HoloLens 2)的多相机系统,从多个视角同步捕获RGB-D视频。
  • 开发了一种多视角位姿优化框架,通过融合多视角间的2D-3D对应关系,提升6DoF位姿估计的精度。
  • 同时使用真实离体手术数据和合成数据进行训练,并通过数据增强提升对遮挡和光照变化的鲁棒性。
  • 应用最先进6DoF位姿估计网络(如BOP基准中的方法),并将其适配于多视角手术器械跟踪场景。
  • 使用真实测试时数据进行域内微调,进一步降低位姿误差并提升泛化能力。
  • 对数据集进行了真实位姿标注,涵盖手术器械、外科医生和患者解剖结构,支持精确评估。

实验结果

研究问题

  • RQ1在真实手术室条件下,多视角RGB-D融合能否实现对手术器械6DoF位姿估计的亚毫米级精度?
  • RQ2不同的相机配置(数量、位置、类型)如何影响位姿精度和对遮挡的鲁棒性?
  • RQ3在真实手术数据上训练时,合成数据在多大程度上能提升模型的泛化能力?
  • RQ4无标记跟踪系统能否在临床相关场景中实现与现有标记导航系统相当或更优的精度?
  • RQ5何种相机配置和数据策略最适于实现实时手术导航中的高精度与高鲁棒性?

主要发现

  • 在最优条件下,使用五台相机的最佳多视角方法对手术钻头实现了1.01 mm的平均位置误差和0.89°的方位误差。
  • 对于螺丝刀,该方法实现了2.79 mm的位置误差和3.33°的方位误差,表明其在不同器械类型下均具有鲁棒性。
  • 即使仅使用两台相机,系统也实现了毫米级精度,证明无标记跟踪可作为标记系统可行的替代方案。
  • 在真实测试时数据上进行域内微调显著降低了位姿误差,凸显了领域特定适应的价值。
  • 合成数据被证明可提升模型鲁棒性,尤其在处理手术室中常见的复杂遮挡和光照变化方面表现突出。
  • 所提出的多视角系统在标记系统因器械离屏而失效时仍能保持功能,展现出更优的遮挡鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。