[Paper Review] Learning Schizophrenia Imaging Genetics Data Via Multiple Kernel Canonical Correlation Analysis
This study proposes a novel framework using Kernel and Multiple Kernel Canonical Correlation Analysis (KCCA/MKCCA) to classify schizophrenia patients from healthy controls by integrating fMRI, DNA methylation, and SNP data. By leveraging nonlinear correlations in high-dimensional imaging genetics data, the method achieves 72.6% classification accuracy—significantly outperforming linear CCA—with the highest accuracy when combining fMRI and DNA methylation data, while SNP data reduces performance.
Kernel and Multiple Kernel Canonical Correlation Analysis (CCA) are employed to classify schizophrenic and healthy patients based on their SNPs, DNA Methylation and fMRI data. Kernel and Multiple Kernel CCA are popular methods for finding nonlinear correlations between high-dimensional datasets. Data was gathered from 183 patients, 79 with schizophrenia and 104 healthy controls. Kernel and Multiple Kernel CCA represent new avenues for studying schizophrenia, because, to our knowledge, these methods have not been used on these data before. Classification is performed via k-means clustering on the kernel matrix outputs of the Kernel and Multiple Kernel CCA algorithm. Accuracies of the Kernel and Multiple Kernel CCA classification are compared to that of the regularized linear CCA algorithm classification, and are found to be significantly more accurate. Both algorithms demonstrate maximal accuracies when the combination of DNA methylation and fMRI data are used, and experience lower accuracies when the SNP data are incorporated.
Motivation & Objective
- To develop a nonlinear multivariate method for integrating multimodal imaging genetics data in schizophrenia research.
- To improve classification accuracy of schizophrenia versus healthy controls beyond traditional linear CCA.
- To investigate the relative contribution of fMRI, DNA methylation, and SNP data to disease classification.
- To evaluate whether nonlinear relationships in high-dimensional omics and neuroimaging data enhance diagnostic discrimination.
- To identify the optimal combination of data modalities for maximal classification performance.
Proposed method
- Kernel and Multiple Kernel CCA are applied to project fMRI, DNA methylation, and SNP data into a Reproducing Kernel Hilbert Space (RKHS) to capture nonlinear correlations.
- The method uses a Gaussian kernel to map input data into a higher-dimensional feature space where nonlinear relationships become linear.
- Canonical correlation coefficients are computed to maximize correlation between projected datasets across multiple modalities.
- k-means clustering is applied to the kernel matrix outputs from KCCA and MKCCA to perform binary classification of patients and controls.
- Regularization is applied to the eigenvalue problem to prevent overfitting and ensure numerical stability.
- The classification performance is evaluated across all combinations of two or three data types, with accuracy measured via cross-validation.
Experimental results
Research questions
- RQ1Can Kernel and Multiple Kernel CCA outperform linear CCA in classifying schizophrenia patients from healthy controls using multimodal imaging genetics data?
- RQ2Which combination of fMRI, DNA methylation, and SNP data yields the highest classification accuracy?
- RQ3Does the inclusion of SNP data degrade classification performance compared to fMRI and DNA methylation alone?
- RQ4Are nonlinear correlations between imaging and epigenetic data more informative for schizophrenia classification than linear correlations?
- RQ5Can the projection coefficients from KCCA/MKCCA serve as effective discriminative features for disease classification?
Key findings
- Kernel and Multiple Kernel CCA achieved a maximum classification accuracy of 72.6%, significantly outperforming linear CCA.
- The highest accuracy was obtained when combining fMRI and DNA methylation data, with no improvement when adding SNP data.
- Classification accuracy was lowest when using only SNP data, indicating limited discriminative power for this modality in this context.
- Both KCCA and MKCCA demonstrated superior performance compared to linear CCA, especially in capturing nonlinear relationships across datasets.
- The combination of fMRI and DNA methylation data produced the most separable components in the kernel space, suggesting stronger biological relevance to schizophrenia.
- The results suggest that epigenetic and neuroimaging data are more informative than genetic variants (SNPs) for classifying schizophrenia in this cohort.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.