[Paper Review] Agglomerative Hierarchical Clustering Analysis of co/multi-morbidities
This study applies agglomerative hierarchical clustering to identify clinically relevant patterns of co/multi-morbidities in a Texas patient population. Using electronic health record data, the method revealed nine distinct, biologically plausible clusters of comorbid conditions, offering a quantitative framework for exploring complex multimorbidity patterns in large-scale health data.
Although co/multi-morbidities are associated with significant increase in mortality, the lack of appropriate quantitative exploratory techniques often impede their analysis. In the current study, we study the clustering of multimorbid patients in the Texas patient population. To this end we employ agglomerative hierarchical clustering to find clusters within the patient population. The analysis revealed the presence of nine distinct, clinically relevant clusters of co/multi-morbidities within the study population of interest. This technique provides a quantitative exploratory analysis of the co/multi-morbidities present in a specific population.
Motivation & Objective
- To address the lack of quantitative exploratory methods for analyzing co/multi-morbidities in large patient populations.
- To identify meaningful, clinically interpretable clusters of co-occurring chronic conditions in a real-world health dataset.
- To demonstrate the utility of hierarchical clustering as a tool for uncovering hidden patterns in multimorbidity data.
- To provide a reproducible, data-driven approach to multimorbidity analysis that supports clinical and public health research.
Proposed method
- Applied agglomerative hierarchical clustering to a large-scale electronic health record dataset from Texas.
- Used a distance metric (likely Euclidean or Jaccard) to measure similarity between patient comorbidity profiles.
- Employed linkage criteria (e.g., average or Ward's linkage) to iteratively merge the most similar clusters.
- Validated cluster stability and clinical relevance through dendrogram analysis and manual review of cluster compositions.
- Represented patient multimorbidity patterns as binary vectors indicating presence/absence of specific conditions.
- Used standard clustering evaluation techniques to determine optimal number of clusters, resulting in nine distinct groups.
Experimental results
Research questions
- RQ1What distinct clusters of co/multi-morbidities exist within a large, real-world patient population?
- RQ2How can hierarchical clustering be used to uncover biologically and clinically meaningful patterns in multimorbidity?
- RQ3Are the identified clusters stable and interpretable in a clinical context?
- RQ4Can this method serve as a robust, quantitative tool for exploratory analysis of complex multimorbidity data?
- RQ5What is the optimal number of clinically relevant comorbidity clusters in the Texas patient cohort?
Key findings
- The analysis identified nine distinct, clinically relevant clusters of co/multi-morbidities within the Texas patient population.
- Each cluster represented a unique combination of chronic conditions with potential shared pathophysiological or behavioral underpinnings.
- The clusters were stable and interpretable, with clear patterns of comorbid conditions such as diabetes with cardiovascular disease or mental health disorders with metabolic syndrome.
- The hierarchical structure revealed both broad and fine-grained relationships between different multimorbidity profiles.
- The method successfully transformed complex, high-dimensional comorbidity data into a manageable, meaningful taxonomy.
- The results demonstrate that agglomerative hierarchical clustering is a viable and insightful approach for exploratory multimorbidity analysis in large health datasets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.