[Paper Review] End-to-End AI-based MRI Reconstruction and Lesion Detection Pipeline for Evaluation of Deep Learning Image Reconstruction
This study presents an end-to-end deep learning pipeline integrating MRI reconstruction and meniscus tear detection using the Gadgetron framework, demonstrating that high-acceleration deep learning reconstructions—despite strong SSIM scores—fail to preserve fine anatomical details critical for lesion detection. The key finding is that reconstruction quality metrics like SSIM do not correlate with clinical performance in pathology detection, highlighting the need for clinically aware evaluation benchmarks.
Deep learning techniques have emerged as a promising approach to highly accelerated MRI. However, recent reconstruction challenges have shown several drawbacks in current deep learning approaches, including the loss of fine image details even using models that perform well in terms of global quality metrics. In this study, we propose an end-to-end deep learning framework for image reconstruction and pathology detection, which enables a clinically aware evaluation of deep learning reconstruction quality. The solution is demonstrated for a use case in detecting meniscal tears on knee MRI studies, ultimately finding a loss of fine image details with common reconstruction methods expressed as a reduced ability to detect important pathology like meniscal tears. Despite the common practice of quantitative reconstruction methodology evaluation with metrics such as SSIM, impaired pathology detection as an automated pathology-based reconstruction evaluation approach suggests existing quantitative methods do not capture clinically important reconstruction outcomes.
Motivation & Objective
- To address the clinical gap in deep learning MRI reconstruction evaluation by incorporating lesion detection as a performance metric.
- To investigate whether high-acceleration deep learning reconstructions preserve fine anatomical details essential for detecting meniscus tears.
- To develop an end-to-end, clinically deployable pipeline using Gadgetron for real-time integration with MRI scanners.
- To compare conventional and deep learning-based reconstruction methods not only by global metrics but also by their impact on automated pathology detection.
- To demonstrate that existing quantitative metrics (e.g., SSIM) may not reflect clinically relevant image quality, especially for subtle pathologies.
Proposed method
- An end-to-end pipeline was built using the Gadgetron framework to process raw k-space data directly from MRI scanners, enabling seamless integration with clinical workflows.
- The pipeline combines multiple MRI reconstruction methods—including zero-filling FFT, CG-SENSE, UNet, and VarNet—at various acceleration rates (R=4 and R=8).
- A YOLO-based object detection network was trained on fully-sampled knee MRI images with bounding box annotations from the fastMRI+ dataset to detect meniscus tears.
- Lesion detection performance was evaluated by comparing true positive (TP) and false negative (FN) detections across different reconstruction methods.
- The framework enables direct comparison of reconstruction algorithms not only via SSIM and NMSE but also through downstream clinical task performance.
- The system was deployed on managed cloud services to support scalable, low-latency testing of reconstruction pipelines.
Experimental results
Research questions
- RQ1Do deep learning MRI reconstruction methods that achieve high SSIM scores still preserve fine anatomical details necessary for detecting meniscus tears?
- RQ2How does the performance of automated meniscus tear detection vary across different reconstruction techniques, including conventional and deep learning-based methods?
- RQ3To what extent do global image quality metrics like SSIM correlate with clinically relevant detection outcomes in MRI reconstruction?
- RQ4Can an end-to-end pipeline integrating reconstruction and lesion detection serve as a robust benchmark for evaluating reconstruction quality beyond traditional metrics?
- RQ5Does training a lesion detection model on reconstructed images (e.g., VarNet) instead of fully-sampled data lead to overfitting or reduced detection accuracy?
Key findings
- At acceleration rate R=8, both UNet and VarNet reconstructions failed to preserve fine meniscus details, resulting in missed detection of meniscus tears despite high SSIM values.
- The YOLO lesion detection network failed to detect meniscus tears in CG-SENSE and zero-filling FFT reconstructions due to low signal-to-noise ratio and artifacts, respectively.
- No significant difference in average SSIM was observed between slices with true positives and false negatives across reconstruction methods, indicating SSIM is not predictive of lesion detection performance.
- Training the YOLO detector on VarNet-reconstructed images (R=4) led to more false positives compared to training on fully-sampled images, suggesting overfitting to reconstruction artifacts.
- The detection network mistakenly identified displaced meniscal tissue as meniscus tears due to visual similarity, indicating limitations in current annotation and model generalization.
- The study demonstrates that existing global metrics like SSIM do not capture clinically important image quality aspects, especially for subtle pathologies, and calls for new pathology-aware evaluation benchmarks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.