Skip to main content
QUICK REVIEW

[논문 리뷰] Interpreting artificial neural networks to detect genome-wide association signals for complex traits

Burak Yelmen, Maris Alver|arXiv (Cornell University)|2024. 07. 26.
Genetic Associations and EpidemiologyBiochemistry, Genetics and Molecular Biology인용 수 3
한 줄 요약

이 연구는 복잡한 형질에서 전장 게놈 연관 분석 신호를 탐지하기 위해 인공 신경망(DNNs)을 해석하기 위한 일반적인 프레임워크를 제안한다. 사후 해석 방법을 활용해 잠재적으로 연관된 유전자좌(PALs)를 식별한다. 에스토니아 생물은행 정신분열증 코hort에 적용한 결과, 높은 정밀도로 PALs를 탐지하였으며, 이는 이전에 보고되지 않은 새로운 유전자좌를 포함하고 있다. 이는 DNNs가 전통적인 선형 GWAS 모델에 대한 실현 가능하고 해석 가능한 대안임을 보여준다.

ABSTRACT

Investigating the genetic architecture of complex diseases is challenging due to the multifactorial and interactive landscape of genomic and environmental influences. Although genome-wide association studies (GWAS) have identified thousands of variants for multiple complex traits, conventional statistical approaches can be limited by simplified assumptions such as linearity and lack of epistasis in models. In this work, we trained artificial neural networks to predict complex traits using both simulated and real genotype-phenotype datasets. We extracted feature importance scores via different post hoc interpretability methods to identify potentially associated loci (PAL) for the target phenotype and devised an approach for obtaining p-values for the detected PAL. Simulations with various parameters demonstrated that associated loci can be detected with good precision using strict selection criteria. By applying our approach to the schizophrenia cohort in the Estonian Biobank, we detected multiple loci associated with this highly polygenic and heritable disorder. There was significant concordance between PAL and loci previously associated with schizophrenia and bipolar disorder, with enrichment analyses of genes within the identified PAL predominantly highlighting terms related to brain morphology and function. With advancements in model optimization and uncertainty quantification, artificial neural networks have the potential to enhance the identification of genomic loci associated with complex diseases, offering a more comprehensive approach for GWAS and serving as initial screening tools for subsequent functional studies.

연구 동기 및 목표

  • 복잡한 형질의 GWAS 맥락에서 딥 네ural 웹(DNNs)을 해석하기 위한 일반적이고 모델에 종속되지 않는 프레임워크를 개발하는 것.
  • 다양한 사후 해석 방법과 로지스틱 회귀 모델 간의 성능을 다양한 시뮬레이션 조건에서 인위적 유전자좌를 탐지하는 데 비교하는 것.
  • 실제 데이터—특히 에스토니아 생물은행 정신분열증 코hort—에 프레임워크를 적용하여 새로운 잠재적으로 연관된 유전자좌(PALs)를 식별하는 것.
  • 해석 도구를 갖춘 DNNs가 전통적인 선형 GWAS 모델의 대안 또는 보완으로서의 유용성을 평가하는 것.
  • 모델 정규화와 확률적 요소가 비선형 유전적 효과(예: 상호작용, 우성) 탐지에 미치는 영향을 평가하는 것.

제안 방법

  • 복잡한 형질을 예측하기 위해 시뮬레이션된 및 실제 유전자형-표현형 데이터셋에 대해 DNNs를 훈련시켰으며, 확률적 요소를 관리하기 위해 강한 드롭아웃과 다수의 랜덤 시드를 사용하였다.
  • LIME, SHAP, 통합 기울기 등 여러 사후 해석 방법을 적용하여 개별 SNP의 중요도 점수를 추출하였다.
  • 엄격한 선택 기준을 사용하여 잠재적으로 연관된 유전자좌(PALs)를 정의하였으며, 중요도 점수와 통계적 임계값을 기반으로 필터링하였다.
  • PALs에 대한 부식 분석을 수행하여 생물학적 관련성을 평가하였으며, 뇌 형태학과 관련된 유전자 영역과 경로에 집중하였다.
  • 다양한 유전적 구조(예: 덧셈 효과, 상호작용 효과, 우성/열성 효과 포함)를 가진 시뮬레이션을 통해 다양한 방법 간 PAL 탐지 성능을 비교하였다.
  • 에스토니아 생물은행(EBB)의 실제 데이터를 사용하여 결과를 검증하였으며, 관련성에 대한 엄격한 품질 기준(관계성 기반 pi-hat < 0.2)과 290,522개의 이형성 SNP로 필터링하였다.
Figure 1: Overview of the approach for obtaining potentially associated loci (PAL) from feature attribution scores obtained via SM, IG and PM methods.
Figure 1: Overview of the approach for obtaining potentially associated loci (PAL) from feature attribution scores obtained via SM, IG and PM methods.

실험 결과

연구 질문

  • RQ1사후 해석 방법을 갖춘 딥 네URAL 웹(DNNs)은 선형 모델보다 비선형 유전적 효과(예: 상호작용, 우성)를 복잡한 형질에서 더 효과적으로 탐지할 수 있는가?
  • RQ2다양한 시뮬레이션 조건에서 다양한 해석 방법(SHAP, LIME, 통합 기울기 등)이 진정한 양성 유전자좌를 식별하는 데 어떻게 비교되는가?
  • RQ3DNN 기반 PAL 탐지 방법은 에스토니아 생물은행 정신분열증 코hort와 같은 실제 생물은행 데이터에서 알려진 또는 새로운 유전자좌를 얼마나 잘 복원할 수 있는가?
  • RQ4모델 정규화와 확률적 요소는 고차원적, 다유전자 형질 데이터에서 PAL 탐지의 신뢰성과 정밀도에 어떻게 영향을 미치는가?
  • RQ5DNNs에 의해 식별된 PALs는 정신분열증과 같은 표적 질환과 관련된 경로(예: 뇌 형태학)에서 생물학적 부식이 유의미하게 나타나는가?

주요 결과

  • 엄격한 선택 기준 하에서 DNN 기반 접근법은 복잡한 유전적 구조를 가진 시뮬레이션에서 높은 정밀도로 연관 유전자좌를 탐지하였다.
  • 로지스틱 회귀 모델과 유사한 ROC AUC를 보였지만, 약간 더 높은 성능를 보였으며, DNNs는 로지스틱 회귀 모델이 탐지하지 못한 유전자좌를 식별하였다. 이는 비선형 효과에 대한 민감도를 시사한다.
  • 모든 해석 방법이 주로 상호작용(상호작용 효과)을 가진 유전자좌를 탐지하였지만, DNNs는 로지스틱 회귀 모델보다 우성 또는 열성 효과를 가진 더 많은 유전자좌를 식별하였다.
  • 유전자 영역에 대한 PALs의 부식 분석 결과, 주로 뇌 형태학과 관련된 용어와 관련된 경로가 우세하게 나타났으며, 이는 정신분열증 코hort에서의 생물학적 관련성을 뒷받침한다.
  • 이 방법은 이전에 문헌에 보고되지 않은 새로운 PALs를 성공적으로 식별하였으며, 기존 GWAS를 초월한 발견 잠재력을 보여주었다.
  • 강한 드롭아웃과 다수의 랜덤 시드를 사용한 앙상블형 훈련이 가짜 양성 결과를 완화시키며, 모델의 확률적 특성에도 불구하고 강건성을 보여주었다.
Figure 2: PAL detected by integrated gradients (IG) approach. Red dashed lines indicate significance thresholds (relaxed and strict) and blue markers indicate PAL above threshold over all trained 10 models (i.e., $PAL_{Common}$ ). For all PAL (a-g), protein coding genes in those regions were provide
Figure 2: PAL detected by integrated gradients (IG) approach. Red dashed lines indicate significance thresholds (relaxed and strict) and blue markers indicate PAL above threshold over all trained 10 models (i.e., $PAL_{Common}$ ). For all PAL (a-g), protein coding genes in those regions were provide

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.