[論文レビュー] Domain-adapted large language models for classifying nuclear medicine reports
本研究では、核医学画像報告書の分類を目的としたドメイン適応型大規模言語モデルの有効性を評価した。具体的には、18F-FDG PET/CT報告書から5段階のデュヴィルスコアを予測することを目的としている。一般向けモデル(例:RoBERTa)をマスク言語モデルを用いて核医学分野特化のテキストで微調整することで、ドメイン適応を実施したところ、5クラス分類の正確度が77.4%に向上した。これは人間の専門医(66%)を上回り、視覚のみのモデルやマルチモーダルモデルをも上回った。
With the growing use of transformer-based language models in medicine, it is unclear how well these models generalize to nuclear medicine which has domain-specific vocabulary and unique reporting styles. In this study, we evaluated the value of domain adaptation in nuclear medicine by adapting language models for the purpose of 5-point Deauville score prediction based on clinical 18F-fluorodeoxyglucose (FDG) PET/CT reports. We retrospectively retrieved 4542 text reports and 1664 images for FDG PET/CT lymphoma exams from 2008-2018 in our clinical imaging database. Deauville scores were removed from the reports and then the remaining text in the reports was used as the model input. Multiple general-purpose transformer language models were used to classify the reports into Deauville scores 1-5. We then adapted the models to the nuclear medicine domain using masked language modeling and assessed its impact on classification performance. The language models were compared against vision models, a multimodal vision language model, and a nuclear medicine physician with seven-fold Monte Carlo cross validation, reported are the mean and standard deviations. Domain adaption improved all language models. For example, BERT improved from 61.3% five-class accuracy to 65.7% following domain adaptation. The best performing model (domain-adapted RoBERTa) achieved a five-class accuracy of 77.4%, which was better than the physician's performance (66%), the best vision model's performance (48.1), and was similar to the multimodal model's performance (77.2). Domain adaptation improved the performance of large language models in interpreting nuclear medicine text reports.
研究の動機と目的
- ドメイン適応が大規模言語モデルの核医学報告書解釈性能を向上させるかどうかを評価すること。
- 自由記述形式のPET/CT報告書に基づいて、リンパ腫の治療反応評価に用いられるデュヴィルスコアを正確に予測できるかどうかを評価すること。
- ドメイン適応型言語モデルの性能を、視覚モデル、マルチモーダルモデル、および人間の専門医と比較すること。
- 核医学分野特化の事前学習が、一般医療分野や一般向けの事前学習(例:BERT)よりも効果的かどうかを特定すること。
- 報告書とスコアが医師によって割り当てられている場合、PET/CTスキャンの画像情報がテキストベース分類に追加的価値をもたらすかどうかを調査すること。
提案手法
- 2008年から2018年までの1つの施設のPACSデータベースから、4,542件のFDG PET/CT報告書と1,664枚の画像を収集し、報告書テキストのN-gram解析によりデュヴィルスコアを抽出した。
- 言語モデルの学習および推論用に、報告書からデュヴィルスコアをマスキングした入力テキストを作成した。
- 核医学分野特化のコーパスを用いて、自己教師付きマスク言語モデルを用いて一般向けトランスフォーマーモデル(例:BERT、RoBERTa)を適応させた。
- 7-foldのモンテカルロ交差検証を用いて、複数の言語モデル、視覚モデル(例:ResNet)、およびマルチモーダル視覚言語モデルを訓練および評価した。
- 主な指標として5クラス分類の正確度を用い、各foldにおける平均±標準偏差を報告した。
- ドメイン適応型モデルを非ドメイン適応型バージョン、視覚モデル、マルチモーダルモデル、および1名の核医学専門医と比較した。
実験結果
リサーチクエスチョン
- RQ1ドメイン適応は、核医学報告書の分類における大規模言語モデルの性能を向上させるか?
- RQ2ドメイン適応型言語モデルの性能は、視覚モデルやマルチモーダルモデルと比較して、同様のタスクで優れているか?
- RQ3ドメイン適応型言語モデルは、テキスト報告書からのデュヴィルスコア予測において、人間の専門医を上回ることができるか?
- RQ4核医学分野特化のドメイン適応は、一般医療分野特化のドメイン適応(例:BioClinicalBERT)や一般向け事前学習(例:BERT)よりも効果的か?
- RQ5報告書とスコアが医師によって割り当てられている場合、PET/CTスキャンの画像情報がテキストベース分類に追加的価値をもたらすか?
主な発見
- ドメイン適応により、全テスト対象言語モデルの性能が著しく向上し、BERTの5クラス分類正確度は61.3%から65.7%に上昇した。
- 最も優れた性能を示したモデル、ドメイン適応型RoBERTaは、5クラス分類正確度77.4% ± 3.4%を達成し、人間の専門医の66%を上回った。
- マルチモーダルモデルも同程度の正確度77.2% ± 3.2%を達成し、予測タスクにおいて言語情報が支配的であることが示された。
- 視覚のみのモデルは性能が低く、最大正確度48.1% ± 3.5%にとどまり、画像特徴のみではこのタスクに不十分であることが示された。
- 核医学テキストで微調整された小型モデルRadBERTは、より大きなRoBERTaモデルと同等の性能を示し、ドメイン特化の事前学習の価値を示した。
- 一般向けおよび医療分野特化の言語モデル(例:BioClinicalBERT)は、ドメイン適応型モデルを上回らなかったため、核医学分野特化の適応が不可欠であることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。