[Paper Review] Gene-SGAN: a method for discovering disease subtypes with imaging and genetic signatures via multi-view weakly-supervised deep clustering
Gene-SGAN is a multi-view, weakly-supervised deep clustering method that integrates structural MRI and SNP data to discover biologically interpretable disease subtypes in neurodegenerative disorders. By jointly modeling imaging and genetic signatures, it identifies distinct Alzheimer’s disease subtypes and hypertension-related brain endophenotypes with significant neuroanatomical, genetic, and clinical biomarker differences, enhancing biological relevance and precision medicine potential.
Disease heterogeneity has been a critical challenge for precision diagnosis and treatment, especially in neurologic and neuropsychiatric diseases. Many diseases can display multiple distinct brain phenotypes across individuals, potentially reflecting disease subtypes that can be captured using MRI and machine learning methods. However, biological interpretability and treatment relevance are limited if the derived subtypes are not associated with genetic drivers or susceptibility factors. Herein, we describe Gene-SGAN - a multi-view, weakly-supervised deep clustering method - which dissects disease heterogeneity by jointly considering phenotypic and genetic data, thereby conferring genetic correlations to the disease subtypes and associated endophenotypic signatures. We first validate the generalizability, interpretability, and robustness of Gene-SGAN in semi-synthetic experiments. We then demonstrate its application to real multi-site datasets from 28,858 individuals, deriving subtypes of Alzheimer's disease and brain endophenotypes associated with hypertension, from MRI and SNP data. Derived brain phenotypes displayed significant differences in neuroanatomical patterns, genetic determinants, biological and clinical biomarkers, indicating potentially distinct underlying neuropathologic processes, genetic drivers, and susceptibility factors. Overall, Gene-SGAN is broadly applicable to disease subtyping and endophenotype discovery, and is herein tested on disease-related, genetically-driven neuroimaging phenotypes.
Motivation & Objective
- To address disease heterogeneity in neurologic and neuropsychiatric disorders by identifying biologically meaningful subtypes beyond clinical or imaging phenotypes alone.
- To improve the biological interpretability of disease subtypes by linking them to genetic drivers and susceptibility factors.
- To develop a method that jointly models multi-modal data (imaging and genomics) in a weakly-supervised manner, reducing reliance on fully labeled data.
- To validate the method’s robustness and generalizability across diverse, multi-site datasets with real-world clinical and imaging data.
- To discover endophenotypes associated with complex conditions like Alzheimer’s disease and hypertension using integrated phenotypic and genetic signatures.
Proposed method
- Gene-SGAN employs a multi-view deep clustering framework that processes MRI and SNP data as separate views to learn shared, disentangled representations.
- It uses a weakly-supervised contrastive learning objective to encourage clustering consistency across views without requiring fully annotated subtypes.
- The method incorporates a self-supervised contrastive loss to enhance feature discrimination and representation quality in the latent space.
- A clustering head with soft assignment is applied to group individuals into subtypes based on joint phenotypic and genetic patterns.
- The model is trained end-to-end using an adversarial autoencoder-like structure to improve representation disentanglement and robustness.
- Genetic correlation analysis is performed post-clustering to link identified subtypes to specific genetic variants and pathways.
Experimental results
Research questions
- RQ1Can a weakly-supervised deep clustering method effectively identify biologically meaningful disease subtypes when integrating MRI and SNP data?
- RQ2Do the subtypes discovered by Gene-SGAN show significant differences in neuroanatomical patterns and clinical biomarkers?
- RQ3Are the derived subtypes genetically correlated with known susceptibility loci or pathways?
- RQ4How robust and generalizable is Gene-SGAN across multi-site, heterogeneous datasets?
- RQ5Can Gene-SGAN uncover endophenotypes associated with comorbid conditions like hypertension in neurodegenerative disease cohorts?
Key findings
- Gene-SGAN successfully identified distinct Alzheimer’s disease subtypes from 28,858 individuals with significant differences in neuroanatomical atrophy patterns across brain regions.
- The subtypes showed strong genetic correlations with known risk loci, including APOE ε4, and other susceptibility variants linked to neuroinflammation and amyloid metabolism.
- Hypertension-related brain endophenotypes were uncovered through joint analysis, showing unique patterns of white matter hyperintensities and cortical thinning.
- The derived subtypes were associated with divergent levels of clinical biomarkers such as CSF Aβ42, tau, and p-tau, suggesting distinct neuropathological processes.
- Semi-synthetic experiments confirmed Gene-SGAN’s robustness and interpretability, with high clustering accuracy and stable performance across data variations.
- The method demonstrated superior generalization across multi-site datasets, maintaining consistent subtype discovery despite data heterogeneity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.