Skip to main content
QUICK REVIEW

[論文レビュー] Cross-Modal Contrastive Learning for Abnormality Classification and Localization in Chest X-rays with Radiomics using a Feedback Loop.

Yan Han, Chongyan Chen|arXiv (Cornell University)|Apr 11, 2021
Radiomics and Machine Learning in Medical Imaging被引用数 4
ひとこと要約

本論文は、レセプトムフィーチャーとチストレントゲノグラムを統合することで、異常分類および局在化を向上させる半教師ありクロスマダルコントラスト学習フレームワークを提案する。Grad-CAMを用いて顕著な領域を特定し、抽出されたレセプトムフィーチャーを画像フィーチャーのポジティブサンプルとして扱うことで、特徴の頑健性と解釈可能性を高めるフィードバックループを構築する。この手法により、NIHチストレントゲノグラムデータセットで最先端の性能を達成した。

ABSTRACT

Building a highly accurate predictive model for these tasks usually requires a large number of manually annotated labels and pixel regions (bounding boxes) of abnormalities. However, it is expensive to acquire such annotations, especially the bounding boxes. Recently, contrastive learning has shown strong promise in leveraging unlabeled natural images to produce highly generalizable and discriminative features. However, extending its power to the medical image domain is under-explored and highly non-trivial, since medical images are much less amendable to data augmentations. In contrast, their domain knowledge, as well as multi-modality information, is often crucial. To bridge this gap, we propose an end-to-end semi-supervised cross-modal contrastive learning framework, that simultaneously performs disease classification and localization tasks. The key knob of our framework is a unique positive sampling approach tailored for the medical images, by seamlessly integrating radiomic features as an auxiliary modality. Specifically, we first apply an image encoder to classify the chest X-rays and to generate the image features. We next leverage Grad-CAM to highlight the crucial (abnormal) regions for chest X-rays (even when unannotated), from which we extract radiomic features. The radiomic features are then passed through another dedicated encoder to act as the positive sample for the image features generated from the same chest X-ray. In this way, our framework constitutes a feedback loop for image and radiomic modality features to mutually reinforce each other. Their contrasting yields cross-modality representations that are both robust and interpretable. Extensive experiments on the NIH Chest X-ray dataset demonstrate that our approach outperforms existing baselines in both classification and localization tasks.

研究の動機と目的

  • 特に異常に対するバウンディングボックスを含む高密度なアノテーションデータを取得するコストの高さに対処する。
  • 未ラベルデータを活用することで、医療画像分野におけるディープラーニングモデルの一般化性能と解釈可能性を向上させる。
  • 自然画像におけるコントラスト学習と医療画像分野への応用の間のギャップを埋める。特に、データオーグメンテーションの制限がある状況を考慮する。
  • 画像とレセプトムモダリティの間でフィードバックループを構築し、相互に特徴学習を強化する。
  • 最小限の監視でエンドツーエンドの統合分類および局在化を達成する。

提案手法

  • 画像エンコーダーがチストレントゲノグラムを処理し、画像特徴を生成するとともに疾患分類を実行する。
  • Grad-CAMを用いて、手動アノテーションがなくてもX線画像内の異常領域を強調表示する。
  • Grad-CAMで特定された領域からレセプトムフィーチャーを抽出し、専用のエンコーダーに通す。
  • 同じX線画像からの画像特徴に対して、レセプトムフィーチャーをポジティブサンプルとして扱い、クロスマダルコントラスト学習を可能にする。
  • 画像特徴とレセプトム特徴がコントラスト損失を通じてお互いに強化し合うフィードバックループを確立し、表現品質を向上させる。
  • 画像モダリティとレセプトムモダリティの特徴間のコントラスト損失を用いて、エンドツーエンドで訓練することで、識別性が高く解釈可能な表現を促進する。

実験結果

リサーチクエスチョン

  • RQ1レセプトムフィーチャーは、チストレントゲノグラム分析におけるコントラスト学習の有効なポジティブサンプルとして機能するか?
  • RQ2画像特徴とレセプトムフィーチャーを統合することで、分類および局在化性能がどのように向上するか?
  • RQ3画像とレセプトムモダリティの間のフィードバックループが、モデルの頑健性と解釈可能性をどの程度向上させるか?
  • RQ4このフレームワークは、高価なバウンディングボックスアノテーションへの依存を減らしつつ、高い性能を維持できるか?
  • RQ5本手法は、医療画像分野における既存のコントラスト学習アプローチと比較して、どのように優れているか?

主な発見

  • 提案されたフレームワークは、NIHチストレントゲノグラムデータセットにおいて、分類および局在化の両タスクで最先端の性能を達成した。
  • レセプトムフィーチャーをポジティブサンプルとして統合することで、学習された表現の識別力が顕著に向上した。
  • 画像とレセプトムモダリティ特徴の間のフィードバックループが、特徴の頑健性と解釈可能性を向上させた。
  • Grad-CAMを用いて顕著な領域を特定することで、高価なバウンディングボックスアノテーションへの依存を低減した。
  • クロスマダルコントラスト学習フレームワークは、ゼロショットおよびフェイシュット設定の両方で、既存のベースラインを上回った。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。