[论文解读] Can artificial intelligence (AI) be used to accurately detect tuberculosis (TB) from chest x-ray? A multiplatform evaluation of five AI products used for TB screening in a high TB-burden setting.
本研究利用来自孟加拉国达卡的大型未见数据集,评估了五种人工智能平台在胸片中检测结核病(TB)的表现。所有人工智能工具的表现均优于人类放射科医生,其中 qXR 的 AUC 达到 0.91,显示出人工智能在高负担环境中进行结核病筛查和分诊的强大潜力。
Powered by artificial intelligence (AI), particularly deep neural networks, computer aided detection (CAD) tools can be trained to recognize TB-related abnormalities on chest radiographs, thereby screening large numbers of people and reducing the pressure on healthcare professionals. Addressing the lack of studies comparing the performance of different products, we evaluated five AI software platforms specific to TB: CAD4TB (v6), InferReadDR (v2), Lunit INSIGHT for Chest Radiography (v4.9.0) , JF CXR-1 (v2) by and qXR (v3) by on an unseen dataset of chest X-rays collected in three TB screening center in Dhaka, Bangladesh. The 23,566 individuals included in the study all received a CXR read by a group of three Bangladeshi board-certified radiologists. A sample of CXRs were re-read by US board-certified radiologists. Xpert was used as the reference standard. All five AI platforms significantly outperformed the human readers. The areas under the receiver operating characteristic curves are qXR: 0.91 (95% CI:0.90-0.91), Lunit INSIGHT CXR: 0.89 (95% CI:0.88-0.89), InferReadDR: 0.85 (95% CI:0.84-0.86), JF CXR-1: 0.85 (95% CI:0.84-0.85), CAD4TB: 0.82 (95% CI:0.81-0.83). We also proposed a new analytical framework that evaluates a screening and triage test and informs threshold selection through tradeoff between cost efficiency and ability to triage. Further, we assessed the performance of the five AI algorithms across the subgroups of age, use cases, and prior TB history, and found that the threshold scores performed differently across different subgroups. The positive results of our evaluation indicate that these AI products can be useful screening and triage tools for active case finding in high TB-burden regions.
研究动机与目标
- 比较五种人工智能平台在高结核病负担地区使用胸片进行结核病筛查的诊断表现。
- 使用标准化参考标准(Xpert)评估人工智能工具相对于人类放射科医生的有效性。
- 开发并应用一种新的分析框架,用于在筛查中平衡成本效率与分诊准确性,以指导阈值选择。
- 评估人工智能模型在不同亚组(如年龄、既往结核病史及临床应用场景)中的表现变异性。
提出的方法
- 在孟加拉国达卡的三个结核病筛查中心收集的 23,566 张胸片数据集中,评估了五种人工智能平台——CAD4TB(v6)、InferReadDR(v2)、Lunit INSIGHT(v4.9.0)、JF CXR-1(v2)和 qXR(v3)的表现。
- 采用 Xpert MTB/RIF 检测作为参考标准以确定结核病阳性状态,并由三位孟加拉国认证的放射科医生进行共识阅片。
- 采用受试者工作特征(ROC)曲线分析计算每种人工智能模型的 AUC 表现。
- 提出一种新颖的分析框架,通过评估成本效率与分诊表现之间的权衡,指导阈值选择。
- 按年龄、临床应用场景及既往结核病史进行亚组分析,以评估模型表现的异质性。
- 通过美国认证放射科医生的重新阅片,对部分胸片进行验证,以确保参考解读的一致性。
实验结果
研究问题
- RQ1与 Xpert 参考标准相比,五种人工智能平台在胸片中检测结核病的表现如何?
- RQ2在高负担、真实世界临床环境中,人工智能模型能否在结核病筛查准确性上超越人类放射科医生?
- RQ3模型表现如何在不同亚组(如年龄、既往结核病史及临床应用场景)中变化?
- RQ4在平衡成本效率与诊断敏感性的情况下,基于人工智能的分诊最优阈值是什么?
- RQ5能否开发一种新的分析框架,以指导人工智能辅助结核病筛查中的阈值选择?
主要发现
- 所有五种人工智能平台在胸片中检测结核病的表现均显著优于人类放射科医生,AUC 范围为 0.82(CAD4TB)至 0.91(qXR)。
- qXR 达到最高诊断表现,AUC 为 0.91(95% CI: 0.90–0.91),其次为 Lunit INSIGHT(AUC 0.89;95% CI: 0.88–0.89)。
- InferReadDR 和 JF CXR-1 的 AUC 均为 0.85(95% CI 分别为 0.84–0.86 和 0.84–0.85)。
- 表现因亚组而异:不同年龄、既往结核病史及临床应用场景下的阈值评分和诊断准确性存在显著差异。
- 所提出的分析框架实现了基于数据的阈值选择,成功平衡了成本效率与分诊表现。
- 本研究证实,基于人工智能的筛查工具可作为高结核病负担地区可靠且可扩展的分诊解决方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。