Skip to main content
QUICK REVIEW

[Paper Review] The Future AI in Healthcare: A Tsunami of False Alarms or a Product of Experts?

Gari D. Clifford|arXiv (Cornell University)|Jul 20, 2020
Healthcare Technology and Patient Monitoring37 references4 citations
TL;DR

The paper proposes a voting ensemble of diverse, independently trained machine learning models as a solution to the overfitting, bias, and poor generalizability plaguing AI in healthcare. By combining algorithms weighted by performance, independence, and contextual features, the approach improves prediction accuracy and provides clinically actionable confidence intervals, with public challenges serving as a scalable mechanism to generate such ensembles and accelerate research.

ABSTRACT

Recent significant increases in affordable and accessible computational power and data storage have enabled machine learning to provide almost unbelievable classification and prediction performances compared to well-trained humans. There have been some promising (but limited) results in the complex healthcare landscape, particularly in imaging. This promise has led some individuals to leap to the conclusion that we will solve an ever-increasing number of problems in human health and medicine by applying `artificial intelligence' to `big (medical) data'. The scientific literature has been inundated with algorithms, outstripping our ability to review them effectively. Unfortunately, I argue that most, if not all of these publications or commercial algorithms make several fundamental errors. I argue that because everyone (and therefore every algorithm) has blind spots, there are multiple `best' algorithms, each of which excels on different types of patients or in different contexts. Consequently, we should vote many algorithms together, weighted by their overall performance, their independence from each other, and a set of features that define the context (i.e., the features that maximally discriminate between the situations when one algorithm outperforms another). This approach not only provides a better performing classifier or predictor but provides confidence intervals so that a clinician can judge how to respond to an alert. Moreover, I argue that a sufficient number of (mostly) independent algorithms that address the same problem can be generated through a large international competition/challenge, lasting many months and define the conditions for a successful event. Finally, I propose introducing the requirement for major grantees to run challenges in the final year of funding to maximize the value of research and select a new generation of grantees.

Motivation & Objective

  • Address the widespread problem of overfitting and poor generalizability in healthcare AI models trained on retrospective medical data.
  • Counter the trend of single-algorithm predictions that lack interpretability and confidence estimation in clinical settings.
  • Propose a framework where multiple diverse algorithms are combined through weighted voting to improve robustness and reliability.
  • Advocate for public competitions (challenges) as a mechanism to generate diverse, high-performing, and independently developed AI models.
  • Reform grant funding by incentivizing challenge-based evaluation and follow-on support for top-performing teams to maximize research impact.

Proposed method

  • Use public, open-access challenges (modeled on PhysioNet) to collect multiple independent machine learning algorithms trained on the same clinical data.
  • Weight each algorithm in the ensemble based on its overall performance, independence from others, and contextual relevance (features that predict when it outperforms others).
  • Apply a voting system where predictions are aggregated using performance-weighted scores to produce a final, more robust classification or prediction.
  • Incorporate confidence intervals into the ensemble output to inform clinicians about prediction reliability and urgency of response.
  • Leverage large-scale, international participation in challenges to generate diverse algorithmic approaches that collectively outperform any single model.
  • Introduce a funding model where major grants require hosting a challenge in the final year, with follow-on grants awarded to top-performing teams.

Experimental results

Research questions

  • RQ1Can a voting ensemble of multiple, independently trained machine learning models outperform any single model in predicting clinical events in healthcare?
  • RQ2How can confidence intervals be meaningfully incorporated into AI predictions to support clinical decision-making?
  • RQ3To what extent can public challenges generate diverse, high-performing, and generalizable AI models for clinical prediction tasks?
  • RQ4What role do training data biases and model development environments play in shaping algorithmic performance and generalizability?
  • RQ5Can challenge-based evaluation replace or supplement traditional peer-review in grant funding to improve research productivity and innovation?

Key findings

  • Most existing AI models in healthcare are overfitted to specific training sets and fail to generalize across different patient populations or clinical contexts.
  • A voting ensemble of multiple diverse algorithms, weighted by performance, independence, and contextual features, significantly improves prediction accuracy and reliability.
  • Public challenges—such as those in the PhysioNet/CinC series—demonstrate that large-scale, international collaboration can generate a sufficient number of independent, high-performing models.
  • Confidence intervals derived from ensemble predictions allow clinicians to assess the reliability of alerts and decide whether to act immediately, delay, or retest.
  • The current grant system often fails to correlate peer-review scores with actual research productivity, suggesting a need for alternative evaluation mechanisms.
  • Requiring challenge-based evaluation and follow-on funding for top-performing teams could maximize research value and accelerate the development of clinically useful AI tools.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.