[Paper Review] From Challenges and Pitfalls to Recommendations and Opportunities: Implementing Federated Learning in Healthcare
This paper reviews federated learning (FL) applications in healthcare up to May 2024, identifying critical methodological flaws—such as poor reproducibility, lack of open-source code, non-IID data, and privacy risks—that hinder clinical utility. It proposes actionable recommendations, including standardized documentation, open model and code sharing, and use of public datasets, to improve FL’s reliability, transparency, and real-world applicability in healthcare.
Federated learning holds great potential for enabling large-scale healthcare research and collaboration across multiple centres while ensuring data privacy and security are not compromised. Although numerous recent studies suggest or utilize federated learning based methods in healthcare, it remains unclear which ones have potential clinical utility. This review paper considers and analyzes the most recent studies up to May 2024 that describe federated learning based methods in healthcare. After a thorough review, we find that the vast majority are not appropriate for clinical use due to their methodological flaws and/or underlying biases which include but are not limited to privacy concerns, generalization issues, and communication costs. As a result, the effectiveness of federated learning in healthcare is significantly compromised. To overcome these challenges, we provide recommendations and promising opportunities that might be implemented to resolve these problems and improve the quality of model development in federated learning with healthcare.
Motivation & Objective
- To evaluate the current state of federated learning (FL) in healthcare, focusing on methodological flaws and clinical applicability.
- To identify recurring challenges such as poor reproducibility, lack of open-source code, and inadequate documentation in FL studies.
- To analyze privacy risks, communication costs, and data heterogeneity (non-IID) as major barriers to clinical deployment.
- To assess the gap between FL research and real-world clinical utility due to insufficient validation and transparency.
- To provide a comprehensive set of recommendations and opportunities to enhance the quality, reproducibility, and clinical relevance of FL in healthcare.
Proposed method
- Conducted a systematic review of recent FL studies in healthcare up to May 2024, focusing on methodological rigor and clinical applicability.
- Evaluated studies based on documentation quality, code availability, use of public datasets, and implementation transparency.
- Identified common issues such as vague terminology (e.g., 'model updates' without specification), custom FL frameworks, and lack of hyperparameter reporting.
- Proposed a standardized FL methodology checklist to improve reproducibility and consistency across studies.
- Advocated for the reuse and extension of established open-source FL frameworks instead of custom implementations to reduce errors and improve community validation.
- Recommended open release of trained models and code to enable independent evaluation and accelerate collaborative progress.
Experimental results
Research questions
- RQ1What are the most prevalent methodological flaws in recent federated learning studies applied to healthcare?
- RQ2Why do most current FL approaches in healthcare fail to achieve clinical utility despite promising performance metrics?
- RQ3How do poor documentation, lack of open-source code, and reliance on private datasets hinder reproducibility and fair evaluation?
- RQ4To what extent do privacy risks, communication costs, and data non-IID characteristics compromise the effectiveness of FL in real-world healthcare settings?
- RQ5What standardized practices and open science principles can improve the reliability, transparency, and clinical translation of FL in healthcare?
Key findings
- Only 27% of reviewed FL studies in healthcare publicly released their code, severely limiting reproducibility and independent validation.
- None of the reviewed studies released their trained models, impeding benchmarking and future model reuse.
- Many studies used custom FL frameworks instead of established open-source tools, increasing the risk of implementation errors and reducing transparency.
- Critical methodological details—such as data preprocessing, augmentation, hyperparameter settings, and definitions of 'model updates'—were frequently missing or ambiguous.
- The reliance on private, non-public datasets hindered fair comparison and generalizability of results across institutions.
- Despite promising performance on centralized benchmarks, most FL models in healthcare lack clinical validation and fail to demonstrate actionable outcomes in real-world settings.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.