Skip to main content
QUICK REVIEW

[论文解读] RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation

MD Wahiduzzaman Khan, Hongwei Sheng|arXiv (Cornell University)|Jul 13, 2023
Retinal Imaging and AnalysisMedicine被引用 3
一句话总结

本论文提出了RVD,这是首个使用手持智能手机采集的基于视频的视网膜血管分割数据集,包含来自415名50至75岁患者的635个视频。该数据集提供多层级空间标注(二值、一般和细粒度动静脉掩码)以及血管搏动的时序标注,揭示了显著的域偏移,对现有模型构成挑战,SVP定位在VTN上的mIoU仅为51.25%,凸显了对新方法的迫切需求。

ABSTRACT

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and the usage of bench-top devices further restricts dataset scalability due to its limited accessibility. Considering these limitations, we introduce the first video-based retinal dataset by employing handheld devices for data acquisition. The dataset comprises 635 smartphone-based fundus videos collected from four different clinics, involving 415 patients from 50 to 75 years old. It delivers comprehensive and precise annotations of retinal structures in both spatial and temporal dimensions, aiming to advance the landscape of vasculature segmentation. Specifically, the dataset provides three levels of spatial annotations: binary vessel masks for overall retinal structure delineation, general vein-artery masks for distinguishing the vein and artery, and fine-grained vein-artery masks for further characterizing the granularities of each artery and vein. In addition, the dataset offers temporal annotations that capture the vessel pulsation characteristics, assisting in detecting ocular diseases that require fine-grained recognition of hemodynamic fluctuation. In application, our dataset exhibits a significant domain shift with respect to data captured by bench-top devices, thus posing great challenges to existing methods. In the experiments, we provide evaluation metrics and benchmark results on our dataset, reflecting both the potential and challenges it offers for vessel segmentation tasks. We hope this challenging dataset would significantly contribute to the development of eye disease diagnosis and early prevention.

研究动机与目标

  • 为解决静态、台式图像数据集在捕捉视网膜血管动态特性(如搏动)方面的局限性。
  • 通过使用便携式智能手机设备进行数据采集,克服传统眼底成像在可扩展性和可及性方面的挑战。
  • 提供大规模、多样化且具有临床相关性的视频数据集,配备全面的空间与时序标注,用于视网膜血管分割。
  • 推动视网膜血管动力学研究,提升糖尿病视网膜病变和青光眼等眼病的早期诊断能力。
  • 为评估视网膜图像分析中的域泛化与模型鲁棒性建立基准。

提出的方法

  • 在四个临床中心使用手持智能手机采集635个眼底视频,确保真实世界中的变异性和可及性。
  • 从每个视频中选取最清晰的帧,生成高质量的二值血管掩码,以表征整体血管结构。
  • 基于血管管径差异对血管进行区分,生成一般动静脉掩码,实现动脉与静脉的识别。
  • 通过将每条血管按宽度划分为子段,生成细粒度动静脉掩码,每个视频产生八个不同的掩码。
  • 通过识别视盘区域血管搏动宽度最大和最小的帧,标注时序动态特性。
  • 建立全面的评估协议,通过跨数据集测试量化RVD与现有图像数据集之间的域差距。

实验结果

研究问题

  • RQ1当在基于视频、由手持设备采集的数据集RVD上进行评估时,最先进视网膜血管分割模型的性能如何泛化?
  • RQ2现有模型在真实世界视频数据中对细微视网膜血管搏动(SVP)的检测与定位失败程度如何?
  • RQ3在强度分布与结构动态特性方面,RVD与传统图像数据集之间存在哪些关键的域差距?
  • RQ4多层级空间标注(二值、一般、细粒度)如何影响血管分割模型的性能与临床实用性?
  • RQ5血管搏动的时序标注能否提升对需要血流动力学分析的眼病的检测能力?

主要发现

  • 在现有图像数据集上训练的模型在RVD上测试时性能显著下降,表明存在显著的域差距。
  • SVP定位任务仍极具挑战性,VTN在RVD上的mIoU仅为51.25%,凸显了对专用时序建模方法的迫切需求。
  • Swin-B-22k在跨数据集评估中于二值分割任务上达到最高的78.67% mIoU,证明大规模预训练的优势。
  • 域差距在强度分布差异中表现明显,如图1(c)所示,这影响了模型的泛化能力。
  • 细粒度动静脉掩码显著增加了标注复杂度与临床相关性,支持对血管病理的精细化分析。
  • 该数据集丰富的时空标注为动态视网膜分析与眼病检测开辟了新的研究方向。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。