[Paper Review] Knowledge-Informed Machine Learning for Cancer Diagnosis and Prognosis: A review
This review surveys knowledge-informed machine learning approaches that integrate biomedical knowledge with data to improve cancer diagnosis and prognosis, covering data types, knowledge representations, and integration strategies.
Cancer remains one of the most challenging diseases to treat in the medical field. Machine learning has enabled in-depth analysis of rich multi-omics profiles and medical imaging for cancer diagnosis and prognosis. Despite these advancements, machine learning models face challenges stemming from limited labeled sample sizes, the intricate interplay of high-dimensionality data types, the inherent heterogeneity observed among patients and within tumors, and concerns about interpretability and consistency with existing biomedical knowledge. One approach to surmount these challenges is to integrate biomedical knowledge into data-driven models, which has proven potential to improve the accuracy, robustness, and interpretability of model results. Here, we review the state-of-the-art machine learning studies that adopted the fusion of biomedical knowledge and data, termed knowledge-informed machine learning, for cancer diagnosis and prognosis. Emphasizing the properties inherent in four primary data types including clinical, imaging, molecular, and treatment data, we highlight modeling considerations relevant to these contexts. We provide an overview of diverse forms of knowledge representation and current strategies of knowledge integration into machine learning pipelines with concrete examples. We conclude the review article by discussing future directions to advance cancer research through knowledge-informed machine learning.
Motivation & Objective
- Motivate the challenge of limited labeled data, high dimensionality, heterogeneity, and interpretability in cancer ML.
- Survey how biomedical knowledge can be integrated into data-driven models to enhance performance and reliability.
- Summarize data types (clinical, imaging, molecular, treatment) and knowledge representations used in cancer ML.
- Discuss current integration strategies and provide concrete examples and future research directions.
Proposed method
- Review state-of-the-art studies that fuse biomedical knowledge with data in cancer ML.
- Categorize knowledge representations and how they are integrated into ML pipelines.
- Highlight modeling considerations for clinical, imaging, molecular, and treatment data contexts.
- Provide concrete examples of knowledge-informed ML approaches in cancer diagnosis and prognosis.
Experimental results
Research questions
- RQ1What forms of biomedical knowledge are used in knowledge-informed ML for cancer problems?
- RQ2How are knowledge representations integrated into ML pipelines across cancer data types (clinical, imaging, molecular, treatment)?
- RQ3What impact does knowledge integration have on accuracy, robustness, and interpretability in cancer diagnosis and prognosis models?
- RQ4What are the current challenges and future directions for knowledge-informed ML in this domain.
Key findings
- Knowledge-informed ML has potential to improve accuracy, robustness, and interpretability of cancer diagnosis and prognosis models.
- Four primary data types are emphasized: clinical, imaging, molecular, and treatment data, each with specific representation and integration considerations.
- The review catalogues diverse knowledge representations and integration strategies with concrete examples.
- The article discusses future directions to advance cancer research through knowledge-informed ML.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.