[論文レビュー] Robust Automatic Whole Brain Extraction on Magnetic Resonance Imaging of Brain Tumor Patients using Dense-Vnet
本稿では、脳腫瘍を有する患者のT1GdおよびFLAIR MRIスキャンから脳組織を正確に分離するための深層学習ベースの全脳抽出手法であるDeepBrainを提案する。自動的に生成された確率マスクで訓練されたモデルは、腫瘍影響下のMRIに対して94.5%のDiceスコア、96.4%の感度、98.5%の特異度を達成し、わずか50例の訓練サンプルでさえも高い耐障害性とデータ効率を示している。
Whole brain extraction, also known as skull stripping, is a process in neuroimaging in which non-brain tissue such as skull, eyeballs, skin, etc. are removed from neuroimages. Skull striping is a preliminary step in presurgical planning, cortical reconstruction, and automatic tumor segmentation. Despite a plethora of skull stripping approaches in the literature, few are sufficiently accurate for processing pathology-presenting MRIs, especially MRIs with brain tumors. In this work we propose a deep learning approach for skull striping common MRI sequences in oncology such as T1-weighted with gadolinium contrast (T1Gd) and T2-weighted fluid attenuated inversion recovery (FLAIR) in patients with brain tumors. We automatically created gray matter, white matter, and CSF probability masks using SPM12 software and merged the masks into one for a final whole-brain mask for model training. Dice agreement, sensitivity, and specificity of the model (referred herein as DeepBrain) was tested against manual brain masks. To assess data efficiency, we retrained our models using progressively fewer training data examples and calculated average dice scores on the test set for the models trained in each round. Further, we tested our model against MRI of healthy brains from the LBP40A dataset. Overall, DeepBrain yielded an average dice score of 94.5%, sensitivity of 96.4%, and specificity of 98.5% on brain tumor data. For healthy brains, model performance improved to a dice score of 96.2%, sensitivity of 96.6% and specificity of 99.2%. The data efficiency experiment showed that, for this specific task, comparable levels of accuracy could have been achieved with as few as 50 training samples. In conclusion, this study demonstrated that a deep learning model trained on minimally processed automatically-generated labels can generate more accurate brain masks on MRI of brain tumor patients within seconds.
研究の動機と目的
- 脳腫瘍患者のMRIスキャンに対して、従来の手法が病理的組織の歪みのため失敗しやすい状況でも、頑健で自動化された全脳抽出手法を開発すること。
- T1GdおよびFLAIRシーケンスに腫瘍を有する場合に、既存の頭蓋骨剥離技術の精度が低いという課題に対処すること。
- 手動セグメンテーションではなく、最小限に処理された自動生成ラベルで訓練された深層学習モデルの性能を評価すること。
- 最小限の訓練サンプルで高い性能を達成できるかを評価することで、データ効率を測定すること。
提案手法
- 本モデルであるDeepBrainは、3D医用画像セグメンテーションを目的としたDense-Vnetアーキテクチャに基づく3D U-Netの変種である。
- SPM12ソフトウェアを用いて灰白質、白質、脳脊髄液の確率マップを生成し、これらを統合して訓練用の単一の全脳マスクとした。
- これらの自動生成ラベルを用いて、脳腫瘍患者のT1GdおよびFLAIR MRIシーケンス上で、モデルをエンドツーエンドで訓練した。
- 手動で描画された脳マスクを基準として、Dice類似係数、感度、特異度を用いてモデルの性能を評価した。
- データ効率の実験として、訓練データの逐次的小さなサブセットでモデルを再訓練し、性能の閾値を評価した。
- 健康な脳MRI(LBP40Aデータセット)を用いた評価を通じて、病態状態に応じた性能差を比較した。
実験結果
リサーチクエスチョン
- RQ1自動生成された確率マスクで訓練された深層学習モデルは、脳腫瘍MRIスキャンにおける全脳抽出で高い精度を達成できるか?
- RQ2本手法で提案されたモデルは、脳腫瘍を有する患者のT1GdおよびFLAIRシーケンスにおいて、手動セグメンテーションのベンチマークと比較してどの程度の性能を示すか?
- RQ3腫瘍影響下のMRIにおける全脳抽出で頑健な性能を達成するために必要な最小訓練サンプル数はどの程度か?
- RQ4本モデルの性能は、脳腫瘍MRIと健康な脳MRIの間でどのように変動するか?
- RQ5限られたデータで訓練された場合、モデルはどの程度の精度を維持できるか。これは、データ効率の程度を示している。
主な発見
- DeepBrainモデルは、脳腫瘍MRIスキャンにおいて平均94.5%のDiceスコアを達成し、手動で描画された脳マスクと高い一致を示した。
- 腫瘍影響下のMRIにおいて、感度96.4%、特異度98.5%が記録され、非脳構造の強力な検出と除外が可能であることが示された。
- LBP40Aデータセットの健康な脳MRIでは、Diceスコアが96.2%に向上し、感度96.6%、特異度99.2%を記録した。
- データ効率の実験により、わずか50例の訓練サンプルでも同等の性能が達成可能であることが示され、限られたデータからの良好な一般化能力が裏付けられた。
- モデルは数秒で正確な脳マスクを生成でき、術前計画や腫瘍セグメンテーションなどの臨床ワークフローへの迅速な応用が可能となった。
- SPM12を用いた自動生成ラベルを訓練に使用することで、時間のかかる手動セグメンテーションを一切不要とし、高い性能を達成できた。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。