Skip to main content
QUICK REVIEW

[論文レビュー] Was there COVID-19 back in 2012? Challenge for AI in Diagnosis with Similar Indications

Imon Banerjee, Priyanshu Sinha|arXiv (Cornell University)|Jun 23, 2020
COVID-19 diagnosis using AI参考文献 19被引用数 5
ひとこと要約

本研究では、外部データセットを用いて、COVID-NetおよびCoroNetの2つの深層学習モデルの一般化性能をCOVID-19の胸部X線画像診断において評価する。内部データでは優れた性能を示すが、両モデルとも顕著な偽陽性率を示しており、特にCOVID-NetはChexPertで55.3%、MIMIC-CXRで23.4%の偽陽性率を示している。これは、過学習およびRT-PCRと画像所見の不一致によるラベルの不均衡や不正確さに起因する一般化性能の欠如を示唆している。

ABSTRACT

Purpose: Since the recent COVID-19 outbreak, there has been an avalanche of research papers applying deep learning based image processing to chest radiographs for detection of the disease. To test the performance of the two top models for CXR COVID-19 diagnosis on external datasets to assess model generalizability. Methods: In this paper, we present our argument regarding the efficiency and applicability of existing deep learning models for COVID-19 diagnosis. We provide results from two popular models - COVID-Net and CoroNet evaluated on three publicly available datasets and an additional institutional dataset collected from EMORY Hospital between January and May 2020, containing patients tested for COVID-19 infection using RT-PCR. Results: There is a large false positive rate (FPR) for COVID-Net on both ChexPert (55.3%) and MIMIC-CXR (23.4%) dataset. On the EMORY Dataset, COVID-Net has 61.4% sensitivity, 0.54 F1-score and 0.49 precision value. The FPR of the CoroNet model is significantly lower across all the datasets as compared to COVID-Net - EMORY(9.1%), ChexPert (1.3%), ChestX-ray14 (0.02%), MIMIC-CXR (0.06%). Conclusion: The models reported good to excellent performance on their internal datasets, however we observed from our testing that their performance dramatically worsened on external data. This is likely from several causes including overfitting models due to lack of appropriate control patients and ground truth labels. The fourth institutional dataset was labeled using RT-PCR, which could be positive without radiographic findings and vice versa. Therefore, a fusion model of both clinical and radiographic data may have better performance and generalization.

研究の動機と目的

  • 内部のCOVID-19データセットで学習された深層学習モデルの外部の現実世界データセットへの一般化性能を評価すること。
  • RT-PCR結果と画像所見の間のラベルの不一致が、モデル性能に与える影響を調査すること。
  • 病院内および公開リポジトリを含む多様なデータセットにおいて、COVID-NetおよびCoroNetという2つの代表的モデルの耐性を比較すること。
  • CX-RベースのCOVID-19診断における現在のAIモデルの限界、特に過学習と高い偽陽性率を特定すること。

提案手法

  • COVID-NetおよびCoroNetを、ChexPert、MIMIC-CXR、ChestX-ray14という3つの公開データセットに加え、2020年1月から5月にかけて収集された病院内EMORY病院データセットで評価した。
  • EMORYデータセットのラベリング基準としてRT-PCR結果をゴールドスタンダードとしたが、画像所見との不一致を認識した。
  • 感度、適合率、F1スコア、および偽陽性率(FPR)という標準指標を用いて、すべてのデータセットで性能を測定した。
  • 異なるデータセット間でのモデル性能を比較し、外部妥当性および一般化能力を評価した。
  • RT-PCR陽性結果と画像所見の異常との不一致を分析し、ラベリングの信頼性を評価した。
  • 臨床的および画像的データの統合が、モデルの耐性および一般化性能の向上に寄与する可能性があると提言した。

実験結果

リサーチクエスチョン

  • RQ1COVID-NetおよびCoroNetは、トレーニング時に使用されなかった外部データセットでどのように性能を発揮するか?
  • RQ2ラベルの混合または不正確さがあるデータセットに適用した際、これらのモデルの偽陽性率はどの程度か?
  • RQ3RT-PCR結果と画像所見の間のラベルの不一致が、モデル性能にどの程度影響を及えるか?
  • RQ4内部データで良好な性能を示すモデルが、現実の臨床データには一般化できないのはなぜか?
  • RQ5臨床的データを画像特徴と統合することで、モデルの一般化性能を向上させ、偽陽性率を低下させることができるか?

主な発見

  • COVID-NetはChexPertデータセットで55.3%の偽陽性率を示し、MIMIC-CXRで23.4%を記録しており、一般化性能が著しく低いことが示された。
  • EMORY病院内データセットでは、COVID-Netは感度61.4%、F1スコア0.54、適合率0.49にとどまり、診断性能が弱いことが明らかになった。
  • CoroNetは顕著に低い偽陽性率を示した:EMORYで9.1%、ChexPertで1.3%、ChestX-ray14で0.02%、MIMIC-CXRで0.06%であった。
  • 両モデルが外部データで性能を低下させたことは、過学習の可能性を示唆しており、これは対照群が不十分で、ラベルノイズがあることが原因である可能性がある。
  • EMORYデータセットにおけるRT-PCR陽性結果と画像所見の不一致は、信頼できるゴールドスタンダードを定義する課題を浮き彫りにした。
  • 本研究は、実臨床環境におけるモデルの耐性および一般化性能を向上させるために、臨床的および画像的データの統合が不可欠であると結論づけた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。