[Paper Review] Unsupervised Extraction of Phenotypes from Cancer Clinical Notes for Association Studies
This paper proposes an unsupervised method to extract phenotypic features from unstructured cancer clinical notes using clustering of medical terms and sentences, enabling association studies with somatic mutation profiles. Applied to 65,000 documents and 3.2 million sentences, the approach identifies 341 significant associations, including 32 novel, biologically plausible hypotheses linking clinical features to genetic mutations.
The recent adoption of Electronic Health Records (EHRs) by health care providers has introduced an important source of data that provides detailed and highly specific insights into patient phenotypes over large cohorts. These datasets, in combination with machine learning and statistical approaches, generate new opportunities for research and clinical care. However, many methods require the patient representations to be in structured formats, while the information in the EHR is often locked in unstructured texts designed for human readability. In this work, we develop the methodology to automatically extract clinical features from clinical narratives from large EHR corpora without the need for prior knowledge. We consider medical terms and sentences appearing in clinical narratives as atomic information units. We propose an efficient clustering strategy suitable for the analysis of large text corpora and to utilize the clusters to represent information about the patient compactly. To demonstrate the utility of our approach, we perform an association study of clinical features with somatic mutation profiles from 4,007 cancer patients and their tumors. We apply the proposed algorithm to a dataset consisting of about 65 thousand documents with a total of about 3.2 million sentences. We identify 341 significant statistical associations between the presence of somatic mutations and clinical features. We annotated these associations according to their novelty, and report several known associations. We also propose 32 testable hypotheses where the underlying biological mechanism does not appear to be known but plausible. These results illustrate that the automated discovery of clinical features is possible and the joint analysis of clinical and genetic datasets can generate appealing new hypotheses.
Motivation & Objective
- To address the challenge of extracting actionable phenotypic information from unstructured clinical narratives in large electronic health record (EHR) datasets.
- To develop a scalable, unsupervised method that does not require prior knowledge or manual annotation of clinical features.
- To enable association studies between extracted clinical features and somatic mutation profiles in cancer patients.
- To discover novel, biologically plausible hypotheses linking clinical phenotypes to tumor genomics.
- To demonstrate the feasibility of automated, large-scale phenome-wide association studies using EHR text.
Proposed method
- The method treats individual medical terms and sentences in clinical notes as atomic information units for analysis.
- It applies an efficient clustering strategy to group similar terms and sentences across a large corpus of 65,000 clinical documents.
- Clusters are used to compactly represent patient-level clinical features, forming structured phenotypic profiles.
- The approach leverages natural language processing and embedding techniques to group semantically similar clinical expressions without supervision.
- Association testing is performed between the presence of each cluster (as a clinical feature) and somatic mutation status across 4,007 cancer patients.
- Statistical significance is assessed, and novel or biologically plausible associations are flagged for further investigation.
Experimental results
Research questions
- RQ1Can unsupervised clustering of clinical text terms and sentences effectively extract meaningful phenotypic features from unstructured EHR notes?
- RQ2What is the extent of overlap between extracted clinical features and known somatic mutation associations?
- RQ3Can the method generate novel, biologically plausible hypotheses linking clinical phenotypes to tumor genomics?
- RQ4How scalable is the approach when applied to large-scale EHR corpora with millions of sentences?
- RQ5To what extent can unsupervised feature extraction support phenome-wide association studies in cancer genomics?
Key findings
- The method successfully extracted 341 statistically significant associations between clinical features and somatic mutations in 4,007 cancer patients.
- Among the 341 associations, 32 were identified as novel and biologically plausible, suggesting potential new links between clinical observations and tumor genomics.
- The approach demonstrated scalability by processing approximately 65,000 clinical documents and 3.2 million sentences.
- Known associations were recovered, validating the method’s ability to detect established clinical-genomic relationships.
- The clustering strategy enabled compact, interpretable representation of patient phenotypes from unstructured text.
- The results show that unsupervised NLP techniques can effectively support large-scale, hypothesis-generating association studies in cancer research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.