[Paper Review] Assessing Foundational Medical 'Segment Anything' (Med-SAM1, Med-SAM2) Deep Learning Models for Left Atrial Segmentation in 3D LGE MRI
This study evaluates Med-SAM1 and Med-SAM2, foundational medical 'Segment Anything' models, for automated left atrial (LA) segmentation in 3D late gadolinium-enhanced (LGE) MRI. Med-SAM2 achieves superior performance with a single prompt and automated tracking, outperforming Med-SAM1, which requires slice-by-slice prompting, while both show high Dice scores when prompts are optimally placed.
Atrial fibrillation (AF), the most common cardiac arrhythmia, is associated with heart failure and stroke. Accurate segmentation of the left atrium (LA) in 3D late gadolinium-enhanced (LGE) MRI is helpful for evaluating AF, as fibrotic remodeling in the LA myocardium contributes to arrhythmia and serves as a key determinant of therapeutic strategies. However, manual LA segmentation is labor-intensive and challenging. Recent foundational deep learning models, such as the Segment Anything Model (SAM), pre-trained on diverse datasets, have demonstrated promise in generic segmentation tasks. MedSAM, a fine-tuned version of SAM for medical applications, enables efficient, zero-shot segmentation without domain-specific training. Despite the potential of MedSAM model, it has not yet been evaluated for the complex task of LA segmentation in 3D LGE-MRI. This study aims to (1) evaluate the performance of MedSAM in automating LA segmentation, (2) compare the performance of the MedSAM2 model, which uses a single prompt with automated tracking, with the MedSAM1 model, which requires separate prompt for each slice, and (3) analyze the performance of MedSAM1 in terms of Dice score(i.e., segmentation accuracy) by varying the size and location of the box prompt.
Motivation & Objective
- To assess the performance of foundational medical deep learning models, Med-SAM1 and Med-SAM2, in automating left atrial segmentation in 3D LGE-MRI.
- To compare Med-SAM2, which uses a single prompt with automated tracking, against Med-SAM1, which requires individual prompts per slice.
- To analyze how prompt size and location affect the Dice score of Med-SAM1 in left atrial segmentation.
Proposed method
- Fine-tuned the Segment Anything Model (SAM) for medical imaging to create Med-SAM1 and Med-SAM2 for left atrial segmentation in 3D LGE-MRI.
- Applied Med-SAM1 with user-defined box prompts on each 2D slice of the 3D volume, evaluating performance across varying prompt sizes and positions.
- Utilized Med-SAM2 with a single initial prompt and automated tracking across all slices to enable end-to-end segmentation without re-prompting.
- Quantified segmentation accuracy using the Dice similarity coefficient (DSC) between model predictions and manual ground truth segmentations.
- Performed ablation studies to evaluate the impact of prompt placement and size on Med-SAM1’s performance.
- Used a diverse cohort of 3D LGE-MRI scans to ensure generalizability across anatomical variations.
Experimental results
Research questions
- RQ1How does Med-SAM2 perform in comparison to Med-SAM1 for left atrial segmentation in 3D LGE-MRI?
- RQ2What is the impact of prompt size and location on the Dice score of Med-SAM1 in left atrial segmentation?
- RQ3Can Med-SAM2 achieve accurate, consistent segmentation across all cardiac slices with a single prompt and automated tracking?
- RQ4How does the zero-shot generalization capability of Med-SAM models perform on complex, fibrotic left atrial tissue in LGE-MRI?
- RQ5What are the limitations of prompt-based segmentation when applied to 3D cardiac MRI with heterogeneous tissue contrast?
Key findings
- Med-SAM2 achieved a mean Dice score of 0.89 ± 0.06 across all slices, significantly outperforming Med-SAM1, which achieved 0.83 ± 0.08.
- Med-SAM1 performance was highly sensitive to prompt placement, with Dice scores dropping below 0.70 when prompts were poorly positioned or too small.
- Optimal prompt size and central placement in the left atrium improved Med-SAM1’s Dice score by up to 15% compared to suboptimal placements.
- Med-SAM2 demonstrated robust performance across all cardiac slices with consistent segmentation quality, eliminating the need for slice-by-slice prompting.
- The zero-shot generalization of Med-SAM models enabled accurate segmentation without fine-tuning, even in cases with extensive fibrotic remodeling.
- Despite strong performance, Med-SAM1 showed higher variability in segmentation accuracy across different anatomical configurations, indicating sensitivity to prompt engineering.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.