Skip to main content
QUICK REVIEW

[論文レビュー] CAP-IQA: Context-Aware Prompt-Guided CT Image Quality Assessment

Kazi Ramisa Rifa, Jie Zhang|arXiv (Cornell University)|Jan 4, 2026
COVID-19 diagnosis using AI被引用数 0
ひとこと要約

CAP-IQA は医療テキスト事前知識とインスタンスレベルの文脈プロンプトおよび因果的デバiasを組み合わせて CT 画像品質を予測し、LDCTIQA 2023 で最先端の相関を達成し、大規模な小児 CT データセットで一般化を示す。

ABSTRACT

Prompt-based methods, which encode medical priors through descriptive text, have been only minimally explored for CT Image Quality Assessment (IQA). While such prompts can embed prior knowledge about diagnostic quality, they often introduce bias by reflecting idealized definitions that may not hold under real-world degradations such as noise, motion artifacts, or scanner variability. To address this, we propose the Context-Aware Prompt-guided Image Quality Assessment (CAP-IQA) framework, which integrates text-level priors with instance-level context prompts and applies causal debiasing to separate idealized knowledge from factual, image-specific degradations. Our framework combines a CNN-based visual encoder with a domain-specific text encoder to assess diagnostic visibility, anatomical clarity, and noise perception in abdominal CT images. The model leverages radiology-style prompts and context-aware fusion to align semantic and perceptual representations. On the 2023 LDCTIQA challenge benchmark, CAP-IQA achieves an overall correlation score of 2.8590 (sum of PLCC, SROCC, and KROCC), surpassing the top-ranked leaderboard team (2.7427) by 4.24%. Moreover, our comprehensive ablation experiments confirm that prompt-guided fusion and the simplified encoder-only design jointly enhance feature alignment and interpretability. Furthermore, evaluation on an in-house dataset of 91,514 pediatric CT images demonstrates the true generalizability of CAP-IQA in assessing perceptual fidelity in a different patient population.

研究の動機と目的

  • CT 画像品質の自動評価を促進し、現実世界の低下劣化下で放射線科医の診断判断を反映する。
  • テキストの医療事前知識と画像特異的文脈プロンプトを組み合わせた CAP-IQA フレームワークを提案する。
  • 因果的デバイアスと動的クロスプロンプト注意機構を用いてプロンプトバイアスを軽減する。
  • LDCTIQA 2023 および社内の小児 CT データセットでの信頼性と一般化能力の優越性を示す。

提案手法

  • 医療事前知識を凍結された PubMedBERT ベースのプロンプト埋め込みでエンコードするテキスト分岐を使用。
  • CT 画像を CNN ベースのエンコーダで処理し、ボトルネック特徴マップとプールされた視覚特徴 f を生成する。
  • f からMLPを介して導出される L 個の画像条件付き文脈プロンプト c′ を導入し、インスタンス適応プロンプト π を形成する。
  • Dynamic Cross-Prompt Attention (DCPA) を適用して視覚特徴とプロンプト特徴を融合し、融合表現を生成する。
  • DCPA 出力をエンコーダ特徴と融合させて CT IQA スコアを [0,4] にスケールして回帰する。
  • 放射線科医に基づくグラウンドトゥルーススコアに対して平均二乗誤差損失で学習する。

実験結果

リサーチクエスチョン

  • RQ1プロンプトで導かれたテキスト由来の医療事前知識は、画像特異的文脈プロンプトと効果的に融合して CT IQA を予測できるか。
  • RQ2動的クロスプロンプト注意が視覚のみまたはテキストのみのベースラインより放射線科医スコアとの整合性を改善するか。
  • RQ3CAP-IQA は施設間・患者集団(例:小児 CT データなど)でどれだけ一般化するのか。

主な発見

  • CAP-IQA は LDCTIQA-test で総合スコア s = 2.8590、相関 r = 0.9866、ρ = 0.9775、τ = 0.8949 を達成。
  • CAP-IQA は LDCTIQA のトップ陣のチーム(s = 2.7427)を 0.1163 上回り、総合相関の改善率は約 4.24% に相当。
  • アブレーション研究により、文脈誘導融合と DyT 正規化が他の選択肢より有効であることを示す。
  • PubMedBERT テキストエンコーダを用いた CNN エンコーダが評価対象アーキテクチャの中で最良の総合性能を示す。
  • 社内の小児 CT データセット(91,514 画像)での評価は、CAP-IQA の異なる集団への一般化性を支持する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。