[Paper Review] Trust and Medical AI: The challenges we face and the expertise needed to overcome them
This paper proposes establishing three specialized expert groups—developers, validators, and operational staff—in digital medicine to address conceptual, technical, and humanistic challenges in medical AI. By creating accredited training programs and governance frameworks, the authors argue that trust in medical AI and healthcare institutions can be preserved through expert-led development, validation, and clinical integration of AI systems.
Artificial intelligence (AI) is increasingly of tremendous interest in the medical field. However, failures of medical AI could have serious consequences for both clinical outcomes and the patient experience. These consequences could erode public trust in AI, which could in turn undermine trust in our healthcare institutions. This article makes two contributions. First, it describes the major conceptual, technical, and humanistic challenges in medical AI. Second, it proposes a solution that hinges on the education and accreditation of new expert groups who specialize in the development, verification, and operation of medical AI technologies. These groups will be required to maintain trust in our healthcare institutions.
Motivation & Objective
- To identify and address the major conceptual, technical, and humanistic challenges hindering the safe and trustworthy deployment of AI in medicine.
- To propose a governance model centered on three expert groups—developers, validators, and operational staff—to ensure responsible AI development and integration in healthcare.
- To advocate for formalized education and accreditation in digital medicine to build a new workforce capable of maintaining public trust in medical AI.
- To prevent erosion of trust in healthcare institutions by proactively managing risks such as bias, adversarial attacks, and data leakage in AI systems.
- To align AI governance in healthcare with the rigor of evidence-based medicine and regulatory standards like those applied to pharmaceuticals.
Proposed method
- Proposes a three-tiered expert framework: developers (designing AI), validators (assessing performance), and operational staff (implementing and monitoring AI in clinical settings).
- Recommends interdisciplinary collaboration between computer scientists and clinicians to ensure medical relevance and technical feasibility of AI applications.
- Advocates for formal accreditation of digital medicine through undergraduate and postgraduate degrees integrating computer science, health science, and medical ethics.
- Suggests adopting evidence-based medicine standards—such as randomized clinical trials for AI model evaluation—beyond predictive accuracy to assess clinical impact.
- Promotes the use of peer review with multidisciplinary critique, including ethical and social implications, to ensure methodological rigor.
- Proposes institutional mechanisms such as 'Turing stamps' and regulatory oversight (e.g., FDA) to formally validate AI systems, mirroring drug safety protocols.
Experimental results
Research questions
- RQ1How can conceptual misunderstandings about AI’s capabilities undermine its safe deployment in clinical settings?
- RQ2What technical challenges arise when applying AI models like LSTMs to complex medical data such as EEG signals, and how can they be mitigated?
- RQ3In what ways can bias, data leakage, and adversarial attacks compromise the reliability and trustworthiness of medical AI systems?
- RQ4How can interdisciplinary collaboration between clinicians and computer scientists improve the design and validation of clinically relevant AI models?
- RQ5What institutional and educational reforms are needed to ensure long-term trust in medical AI through expert governance and clinical integration?
Key findings
- Failure to clearly define research questions and hypotheses in AI studies can lead to flawed model design, such as training on inpatient data while intending to deploy as a community screening tool.
- Overfitting and data leakage are significant technical risks that can invalidate model performance if not properly understood and mitigated by analysts.
- AI systems may inadvertently replicate biases present in training data, especially in conditions without consensus in clinical nosology or pathophysiology.
- Operational staff often reject AI recommendations due to obscurity or lack of clinical relevance, highlighting the need for human-centered design and clinician literacy in AI.
- Formal validation of AI in healthcare should extend beyond predictive accuracy to include clinical outcomes, using standards comparable to those in evidence-based medicine.
- The creation of accredited digital medicine programs and specialized roles such as 'digital doctors' and 'digital nurses' is essential for maintaining patient trust and ensuring responsible AI deployment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.