Skip to main content
QUICK REVIEW

[论文解读] A clinical validation of VinDr-CXR, an AI system for detecting abnormal chest radiographs

Ngoc Huy Nguyen, Ha Q. Nguyen|arXiv (Cornell University)|Apr 6, 2021
COVID-19 diagnosis using AI参考文献 24被引用 6
一句话总结

本研究通过将AI系统VinDr-CXR整合到越南一家真实省级医院的PACS系统中,前瞻性地验证了其在检测胸部X光片异常方面的性能。该系统在将AI预测结果与医院HIS系统中提取的放射科报告进行对比时,取得了0.653的F1分数(95%置信区间为0.635–0.671),尽管与实验室环境下的结果相比有所下降,但其在真实临床环境中的表现依然稳健。

ABSTRACT

Computer-Aided Diagnosis (CAD) systems for chest radiographs using artificial intelligence (AI) have recently shown a great potential as a second opinion for radiologists. The performances of such systems, however, were mostly evaluated on a fixed dataset in a retrospective manner and, thus, far from the real performances in clinical practice. In this work, we demonstrate a mechanism for validating an AI-based system for detecting abnormalities on X-ray scans, VinDr-CXR, at the Phu Tho General Hospital - a provincial hospital in the North of Vietnam. The AI system was directly integrated into the Picture Archiving and Communication System (PACS) of the hospital after being trained on a fixed annotated dataset from other sources. The performance of the system was prospectively measured by matching and comparing the AI results with the radiology reports of 6,285 chest X-ray examinations extracted from the Hospital Information System (HIS) over the last two months of 2020. The normal/abnormal status of a radiology report was determined by a set of rules and served as the ground truth. Our system achieves an F1 score - the harmonic average of the recall and the precision - of 0.653 (95% CI 0.635, 0.671) for detecting any abnormalities on chest X-rays. Despite a significant drop from the in-lab performance, this result establishes a high level of confidence in applying such a system in real-life situations.

研究动机与目标

  • 评估基于AI的CAD系统在临床环境中进行胸部X光片解读的实际表现。
  • 开发并应用一种方法,利用现有医院信息系统对AI系统进行前瞻性验证,而无需直接集成PACS和HIS系统。
  • 通过在放射科报告中应用基于模板的规则,建立正常/异常分类的可靠真实情况(ground truth)。
  • 评估VinDr-CXR在真实世界、非实验室环境中,面对多样化且未经筛选的临床数据时的泛化能力。
  • 为未来医学影像AI系统在临床环境中的验证提供基准参考。

提出的方法

  • 将VinDr-CXR系统直接集成到越南富寿省综合医院的影像归档与通信系统(PACS)中。
  • 系统在为期两个月(2020年11月至12月)内前瞻性处理了6,687例胸部X光检查,为每项检查生成AI预测结果。
  • 使用XML解析器从医院信息系统(HIS)中提取放射科报告,形成包含6,687份报告的数据集。
  • 通过患者ID和DICOM属性,使用匹配算法将AI结果与对应的放射科报告关联,最终获得6,285份匹配的检查。
  • 对放射科报告应用模板匹配规则,基于四个解剖区域(胸壁、胸膜、肺部、纵隔)中的预定义短语,将报告分类为正常或异常。
  • 通过10,000次自举重采样计算F1分数,并基于自举分布得出95%置信区间,以评估性能。

实验结果

研究问题

  • RQ1在未重新训练的情况下,基于AI的CAD系统在真实临床环境中能否保持可靠的性能?
  • RQ2在前瞻性真实世界部署中,基于回顾性数据集训练的AI系统性能与在实验室环境下的表现相比如何?
  • RQ3在临床验证环境中,放射科报告模板在建立正常/异常分类的可靠真实情况方面,其适用程度如何?
  • RQ4数据分布变化和临床背景对放射科AI模型在真实世界中的表现有何影响?

主要发现

  • 基于6,285份匹配检查,VinDr-CXR系统在真实医院环境中检测胸部X光片异常的F1得分为0.653(95%置信区间为0.635–0.671)。
  • 该F1得分相比实验室环境下的0.831有显著下降,表明性能差距可能源于数据分布变化或临床背景的影响。
  • 真实情况通过在放射科报告中应用模板匹配建立,结果为4,529例(72.4%)正常和1,756例(27.6%)异常。
  • 使用F1得分的自举分布计算了95%置信区间,确保了性能估计的稳健性。
  • 尽管性能有所下降,0.653的F1得分仍为临床部署提供了高度信心,优于先前研究中报告的0.435的肺炎检测模型F1得分。
  • 本研究展示了利用现有医院IT基础设施对医学影像AI系统进行前瞻性、真实世界验证的可行框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。