Skip to main content
QUICK REVIEW

[论文解读] Cross-trait prediction accuracy of high-dimensional ridge-type estimators in genome-wide association studies

Bingxin Zhao, Hongtu Zhu|arXiv (Cornell University)|Nov 22, 2019
Genetic Associations and Epidemiology参考文献 95被引用 6
一句话总结

本文研究了在具有密集遗传结构的高维全基因组关联研究(GWAS)中,岭型估计量(包括边际估计量和最佳线性无偏预测BLUP)的预测精度。结果表明,当特征数与样本量之比 ω = p/n 超过 5 时,样本外 R² 在所有正则化参数 λ 下几乎达到最优,表明在高维情形下,λ 的选择对预测精度的影响可忽略不计。

ABSTRACT

Marginal association summary statistics have attracted great attention in statistical genetics, mainly because the primary results of most genome-wide association studies (GWAS) are produced by marginal screening. In this paper, we study the prediction accuracy of marginal estimator in dense (or sparsity free) high-dimensional settings with $(n,p,m) o \infty$, $m/n o γ\in (0,\infty)$, and $p/n o ω\in (0,\infty)$. We consider a general correlation structure among the $p$ features and allow an unknown subset $m$ of them to be signals. As the marginal estimator can be viewed as a ridge estimator with regularization parameter $λ o \infty$, we further investigate a class of ridge-type estimators in a unifying framework, including the popular best linear unbiased prediction (BLUP) in genetics. We find that the influence of $λ$ on out-of-sample prediction accuracy heavily depends on $ω$. Though selecting an optimal $λ$ can be important when $p$ and $n$ are comparable, it turns out that the out-of-sample $R^2$ of ridge-type estimators becomes near-optimal for any $λ\in (0,\infty)$ as $ω$ increases. For example, when features are independent, the out-of-sample $R^2$ is always bounded by $1/ω$ from above and is largely invariant to $λ$ given large $ω$ (say, $ω>5$). We also find that in-sample $R^2$ has completely different patterns and depends much more on $λ$ than out-of-sample $R^2$. In practice, our analysis delivers useful messages for genome-wide polygenic risk prediction and computation-accuracy trade-off in dense high-dimensions. We numerically illustrate our results in simulation studies and a real data example.

研究动机与目标

  • 评估在 (n,p,m) → ∞ 的密集高维GWAS设定下,岭型估计量的样本外预测精度。
  • 理解当 p 和 n 较大且可比时,正则化参数 λ 如何影响预测性能。
  • 比较边际估计量(λ → ∞)与岭估计量和BLUP估计量在预测精度和样本内拟合方面的表现。
  • 考察特征数与样本量之比 ω = p/n 在决定预测精度对 λ 敏感性方面的作用。
  • 为在高维设定下基于汇总统计量构建多基因风险评分提供理论与实证指导。

提出的方法

  • 构建一个包含 p 个特征、m 个信号和 n 个样本的高维线性模型,假设SNP之间存在一般相关结构 Σ。
  • 将边际估计量视为 λ → ∞ 时的岭估计量,并研究包含BLUP在内的统一岭型估计量类。
  • 在 n,p,m → ∞ 且 m/n → γ、p/n → ω 的极限下,推导样本外 R² 和样本内 R² 的渐近表达式。
  • 利用随机矩阵理论和渐近分析,刻画预测精度随 ω 和 λ 变化的行为。
  • 通过模拟研究和一个真实神经影像GWAS数据集(PING队列)验证理论发现。
  • 在独立和分块对角相关结构下,比较边际、岭和BLUP估计量的预测性能。

实验结果

研究问题

  • RQ1在高维GWAS中,岭型估计量的样本外预测精度如何依赖于正则化参数 λ?
  • RQ2特征数与样本量之比 ω = p/n 对预测精度对 λ 的敏感性有何影响?
  • RQ3边际估计量(λ → ∞)的预测精度与最优岭估计量和BLUP估计量相比如何?
  • RQ4为何在高维设定下,样本内 R² 比样本外 R² 更敏感于 λ?
  • RQ5在何种条件下样本外 R² 对 λ 不再敏感?这对多基因风险评分建模有何启示?

主要发现

  • 当 ω = p/n > 5 时,岭型估计量的样本外 R² 几乎达到最优且对 λ 不敏感,无论真实信号比例如何。
  • 在特征独立时,样本外 R² 的上界为 1/ω,且随着 ω 增大而趋近该上界。
  • 当 ω > 5 时,边际估计量(λ → ∞)可实现接近最优的样本外 R²,表明在此类情形下 λ 选择不再必要。
  • 相比之下,样本内 R² 对 λ 依赖性强,凸显了样本内与样本外性能之间的根本不对称性。
  • 在高维设定下(ω 较大),最优 λ 对预测的影响较小,从而在实践中降低了计算负担。
  • 模拟和 PING 队列的实证结果表明,当 ω > 5 时,预测精度在不同 λ 下趋于稳定,支持基于汇总统计量的多基因风险预测的稳健性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。