Skip to main content
QUICK REVIEW

[论文解读] Interpreting artificial neural networks to detect genome-wide association signals for complex traits

Burak Yelmen, Maris Alver|arXiv (Cornell University)|Jul 26, 2024
Genetic Associations and EpidemiologyBiochemistry, Genetics and Molecular Biology被引用 3
一句话总结

本研究提出一种通用框架,用于通过事后可解释性方法解释深度神经网络(DNNs),以在复杂性状中检测全基因组关联信号,识别潜在关联位点(PALs)。该方法应用于爱沙尼亚生物样本库的精神分裂症队列,以高精度检测到PALs,包括此前未报道的新位点,证明DNNs是传统线性GWAS模型的可行且可解释的替代方案。

ABSTRACT

Investigating the genetic architecture of complex diseases is challenging due to the multifactorial and interactive landscape of genomic and environmental influences. Although genome-wide association studies (GWAS) have identified thousands of variants for multiple complex traits, conventional statistical approaches can be limited by simplified assumptions such as linearity and lack of epistasis in models. In this work, we trained artificial neural networks to predict complex traits using both simulated and real genotype-phenotype datasets. We extracted feature importance scores via different post hoc interpretability methods to identify potentially associated loci (PAL) for the target phenotype and devised an approach for obtaining p-values for the detected PAL. Simulations with various parameters demonstrated that associated loci can be detected with good precision using strict selection criteria. By applying our approach to the schizophrenia cohort in the Estonian Biobank, we detected multiple loci associated with this highly polygenic and heritable disorder. There was significant concordance between PAL and loci previously associated with schizophrenia and bipolar disorder, with enrichment analyses of genes within the identified PAL predominantly highlighting terms related to brain morphology and function. With advancements in model optimization and uncertainty quantification, artificial neural networks have the potential to enhance the identification of genomic loci associated with complex diseases, offering a more comprehensive approach for GWAS and serving as initial screening tools for subsequent functional studies.

研究动机与目标

  • 开发一种通用的、与模型无关的框架,用于在复杂性状的GWAS背景下解释深度神经网络。
  • 在不同模拟条件下,比较多种事后可解释性方法与逻辑回归在检测致病位点方面的性能。
  • 将该框架应用于真实世界数据——特别是爱沙尼亚生物样本库的精神分裂症队列——以识别新的潜在关联位点(PALs)。
  • 评估结合可解释性工具的DNN在作为传统线性GWAS模型的替代或补充方面的实用性。
  • 评估模型正则化和随机性对检测非线性遗传效应(如上位性、显性)的影响。

提出的方法

  • 在模拟和真实基因型-表型数据集上训练深度神经网络(DNNs),以预测复杂性状,采用高比例丢弃率和多个随机种子以控制随机性。
  • 应用多种事后可解释性方法(包括LIME、SHAP和积分梯度)提取单个SNP的特征重要性得分。
  • 采用严格的选择标准定义潜在关联位点(PALs),基于重要性得分和统计阈值进行过滤。
  • 对PALs进行富集分析,以评估其生物相关性,重点关注与脑形态相关的基因区域和通路。
  • 通过包含加性效应、上位性效应以及显性/隐性效应的模拟,比较不同方法在PAL检测性能上的表现。
  • 使用来自爱沙尼亚生物样本库(EstBB)的真实世界数据进行验证,严格质量控制包括相关性pi-hat < 0.2,以及筛选至290,522个双等位基因SNP。
Figure 1: Overview of the approach for obtaining potentially associated loci (PAL) from feature attribution scores obtained via SM, IG and PM methods.
Figure 1: Overview of the approach for obtaining potentially associated loci (PAL) from feature attribution scores obtained via SM, IG and PM methods.

实验结果

研究问题

  • RQ1基于事后可解释性方法的深度神经网络是否能比线性模型更有效地检测复杂性状中的非线性遗传效应(如上位性、显性)?
  • RQ2在不同模拟参数下,不同可解释性方法(如SHAP、LIME、积分梯度)在识别真正阳性位点方面的表现如何比较?
  • RQ3在真实世界生物样本库数据(如爱沙尼亚生物样本库精神分裂症队列)中,基于DNN的PAL检测在多大程度上能恢复已知或新发现的位点?
  • RQ4模型正则化和随机性如何影响在高维、多基因数据中PAL检测的可靠性与精确度?
  • RQ5由DNN识别的PALs是否在与目标疾病(如精神分裂症中的脑形态)相关的通路上表现出显著的生物富集?

主要发现

  • 在严格选择标准下,DNN方法在复杂遗传架构的模拟中以高精度检测到关联位点。
  • 尽管曲线下面积(ROC AUC)与逻辑回归相当(略优),但DNN识别出的位点与逻辑回归未检测到的位点不同,表明其对非线性效应更敏感。
  • 所有可解释性方法主要检测到具有交互作用(上位性)效应的位点,但DNN识别出的显性或隐性效应位点数量多于逻辑回归。
  • 对PALs在基因区域的富集分析主要与脑形态相关术语相关,支持其在精神分裂症队列中的生物学相关性。
  • 该方法成功识别出文献中此前未报道的新PALs,表明其在发现超越传统GWAS的位点方面具有潜力。
  • 采用高比例丢弃率和多随机种子的集成式训练策略有助于减少假阳性,表明即使在模型随机性下仍具稳健性。
Figure 2: PAL detected by integrated gradients (IG) approach. Red dashed lines indicate significance thresholds (relaxed and strict) and blue markers indicate PAL above threshold over all trained 10 models (i.e., $PAL_{Common}$ ). For all PAL (a-g), protein coding genes in those regions were provide
Figure 2: PAL detected by integrated gradients (IG) approach. Red dashed lines indicate significance thresholds (relaxed and strict) and blue markers indicate PAL above threshold over all trained 10 models (i.e., $PAL_{Common}$ ). For all PAL (a-g), protein coding genes in those regions were provide

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。