Skip to main content
QUICK REVIEW

[論文レビュー] Semi-parametric Bayesian variable selection for gene-environment interactions

Jie Ren, Fei Zhou|arXiv (Cornell University)|Jun 3, 2019
Genetic Associations and Epidemiology参考文献 44被引用数 5
ひとこと要約

本稿では、高次元ゲノムデータにおいて線形および非線形遺伝子×環境(G×E)相互作用を同時に同定する、新しい半パラメトリックベイズ変数選択モデルを提案する。非線形効果をモデル化するためにBスプラインを用い、個々の要因およびグループレベルで階層的スパイクアンドスラブ事前分布を適用することで、自動的な構造的同定が可能となる—主効果のみ、G×E相互作用、遺伝的効果なしの3つの状態を区別できる。シミュレーションおよび実データ解析において優れた性能を示した。

ABSTRACT

Many complex diseases are known to be affected by the interactions between genetic variants and environmental exposures beyond the main genetic and environmental effects. Study of gene-environment (G$ imes$E) interactions is important for elucidating the disease etiology. Existing Bayesian methods for G$ imes$E interaction studies are challenged by the high-dimensional nature of the study and the complexity of environmental influences. Many studies have shown the advantages of penalization methods in detecting G$ imes$E interactions in "large p, small n" settings. However, Bayesian variable selection, which can provide fresh insight into G$ imes$E study, has not been widely examined. We propose a novel and powerful semi-parametric Bayesian variable selection model that can investigate linear and nonlinear G$ imes$E interactions simultaneously. Furthermore, the proposed method can conduct structural identification by distinguishing nonlinear interactions from main-effects-only case within the Bayesian framework. Spike and slab priors are incorporated on both individual and group levels to identify the sparse main and interaction effects. The proposed method conducts Bayesian variable selection more efficiently than existing methods. Simulation shows that the proposed model outperforms competing alternatives in terms of both identification and prediction. The proposed Bayesian method leads to the identification of main and interaction effects with important implications in a high-throughput profiling study with high-dimensional SNP data.

研究の動機と目的

  • 「大p、小n」設定における、従来のベイズ手法が非線形G×E相互作用を検出する際の限界を克服すること。
  • 線形および非線形G×E相互作用を同時に同定できる手法の開発。
  • ベイズ枠組み内で、主効果のみ、G×E相互作用、遺伝的効果なしの3つの状態を区別する構造的同定を可能にすること。
  • 従来のペナルティ法およびベイズ的手法と比較して、変数選択の効率性および予測精度を向上させること。

提案手法

  • 非線形G×E相互作用を柔軟に捉えるために、Bスプライン基底展開を用いてモデル化する。
  • 個々の要因およびグループレベルでスパイクアンドスラブ事前分布を適用した階層的ベイズモデルを実装し、重要な効果のスパース選択を可能にする。
  • グループごとのスプライン係数に多変量ラプラス事前分布、個々の係数に単変量ラプラス事前分布を適用し、縮小と選択を誘導する。
  • C++で加速されたコアモジュールを備えた効率的なギブスサンプラーを用い、高次元設定における高速なMCMC計算を実現する。
  • スプライン基底変換を用いて、係数関数の変動、非ゼロ定数、ゼロの3種類を区別することで、自動的な構造的同定を実現する。
  • 臨床的共変量を組み込み、交絡要因をモデルフレームワーク内で適切に扱うことで、推定精度を向上させる。

実験結果

リサーチクエスチョン

  • RQ1ベイズ変数選択手法は、高次元ゲノムデータにおける線形および非線形遺伝子×環境相互作用を効果的に検出できるか?
  • RQ2提案手法は、統合フレームワーク内で、主効果のみ、G×E相互作用、遺伝的効果なしの3状態をどれほど正確に区別できるか?
  • RQ3Bスプラインを用いた半パラメトリックモデリングの導入により、パラメトリック代替手法と比較して検出力および予測精度が向上するか?
  • RQ4階層的スパイクアンドスラブ事前分布構造は、高次元G×E相互作用研究における変数選択効率をどの程度向上させるか?
  • RQ5実際の高スループットSNPデータにおいて、従来のペナルティ法およびベイズ手法と比較して、提案手法はどの程度の性能を示すか?

主な発見

  • 提案手法は、複数のシミュレーション設定において、競合手法と比較して変数同定および予測精度の両面で優れた性能を示した。
  • 看護師健康調査データにおいて、非線形G×E相互作用が有意に検出された。図1は、線形相互作用仮定の明確な違反を示している。
  • 実データ応用において、rs1106380、rs10999234、rs796945などのSNPを含む有意なG×E相互作用が同定され、信頼区間が生物学的妥当性を支持した。
  • 階層的スパイクアンドスラブ事前分布構造により、個々の要因およびグループレベルで重要な効果の効率的選択が可能となり、誤検出の低減が達成された。
  • MCMCアルゴリズムは高速な収束と計算を達成し、コアモジュールはC++で実装されており、スケーラビリティに優れた性能を発揮した。
  • 主効果のみと真のG×E相互作用の区別において、本手法は頑健に機能し、構造的同定能力を確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。