[Paper Review] Interpreting artificial neural networks to detect genome-wide association signals for complex traits
This study proposes a general framework to interpret artificial neural networks (DNNs) for detecting genome-wide association signals in complex traits, using post-hoc interpretability methods to identify potentially associated loci (PALs). Applied to the Estonian Biobank schizophrenia cohort, the method detected PALs with high precision, including novel loci not previously reported, demonstrating DNNs as a viable, interpretable alternative to conventional linear GWAS models.
Investigating the genetic architecture of complex diseases is challenging due to the multifactorial and interactive landscape of genomic and environmental influences. Although genome-wide association studies (GWAS) have identified thousands of variants for multiple complex traits, conventional statistical approaches can be limited by simplified assumptions such as linearity and lack of epistasis in models. In this work, we trained artificial neural networks to predict complex traits using both simulated and real genotype-phenotype datasets. We extracted feature importance scores via different post hoc interpretability methods to identify potentially associated loci (PAL) for the target phenotype and devised an approach for obtaining p-values for the detected PAL. Simulations with various parameters demonstrated that associated loci can be detected with good precision using strict selection criteria. By applying our approach to the schizophrenia cohort in the Estonian Biobank, we detected multiple loci associated with this highly polygenic and heritable disorder. There was significant concordance between PAL and loci previously associated with schizophrenia and bipolar disorder, with enrichment analyses of genes within the identified PAL predominantly highlighting terms related to brain morphology and function. With advancements in model optimization and uncertainty quantification, artificial neural networks have the potential to enhance the identification of genomic loci associated with complex diseases, offering a more comprehensive approach for GWAS and serving as initial screening tools for subsequent functional studies.
Motivation & Objective
- To develop a general, model-agnostic framework for interpreting deep neural networks in the context of GWAS for complex traits.
- To compare the performance of various post-hoc interpretability methods and logistic regression in detecting causal loci under diverse simulation conditions.
- To apply the framework to real-world data—specifically the Estonian Biobank schizophrenia cohort—to identify novel potentially associated loci (PALs).
- To evaluate the utility of DNNs with interpretability tools as an alternative or complement to conventional linear GWAS models.
- To assess the impact of model regularization and stochasticity on detecting non-linear genetic effects such as epistasis and dominance.
Proposed method
- Trained deep neural networks (DNNs) on simulated and real genotype-phenotype datasets to predict complex traits, using heavy dropout and multiple random seeds to manage stochasticity.
- Applied multiple post-hoc interpretability methods—including LIME, SHAP, and integrated gradients—to extract feature importance scores for individual SNPs.
- Used strict selection criteria to define potentially associated loci (PALs), filtering based on importance scores and statistical thresholds.
- Performed enrichment analyses on PALs to assess biological relevance, focusing on genic regions and pathways related to brain morphology.
- Compared PAL detection performance across methods using simulations with varying genetic architectures, including additive, epistatic, and dominant/recessive effects.
- Validated results using real-world data from the Estonian Biobank (EstBB), with strict quality control including pi-hat < 0.2 for relatedness and filtering to 290,522 bi-allelic SNPs.

Experimental results
Research questions
- RQ1Can deep neural networks with post-hoc interpretability methods detect non-linear genetic effects (e.g., epistasis, dominance) in complex traits more effectively than linear models?
- RQ2How do different interpretability methods (e.g., SHAP, LIME, integrated gradients) compare in identifying true positive loci under varying simulation parameters?
- RQ3To what extent can DNN-based PAL detection recover known or novel loci in real-world biobank data, such as the Estonian Biobank schizophrenia cohort?
- RQ4How does model regularization and stochasticity affect the reliability and precision of PAL detection in high-dimensional, polygenic data?
- RQ5Do PALs identified by DNNs show significant biological enrichment in pathways relevant to the target disease, such as brain morphology in schizophrenia?
Key findings
- The DNN-based approach detected associated loci with high precision under strict selection criteria, particularly in simulations with complex genetic architectures.
- Despite comparable ROC AUC to logistic regression (marginally better for LR), the DNNs identified distinct loci not detected by logistic regression, suggesting sensitivity to non-linear effects.
- All interpretability methods primarily detected loci with interactive (epistatic) effects, but DNNs identified a higher number of loci with dominant or recessive effects compared to logistic regression.
- Enrichment analyses of PALs in genic regions were predominantly associated with terms related to brain morphology, supporting biological relevance in the schizophrenia cohort.
- The method successfully identified novel PALs not previously reported in the literature, indicating potential for discovery beyond conventional GWAS.
- Heavy dropout and ensemble-like training with multiple random seeds helped mitigate false positives, suggesting robustness despite model stochasticity.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.