[论文解读] On Genetic Correlation Estimation With Summary Statistics From Genome-Wide Association Studies
本文识别并纠正了在高度多基因或全基因组特征中交叉特征多基因风险评分(PRS)估计中的根本性渐近偏差,该偏差系统性地低估了遗传相关性。作者提出了一种一致的PRS估计量,以及一种仅基于GWAS汇总统计量的新遗传相关性估计量,证明该偏差与因果SNP数量无关,而是由三重组合(n, p, h²)驱动,并通过英国生物银行数据进行了有限样本验证。
Cross-trait polygenic risk score (PRS) method has gained popularity for assessing genetic correlation of complex traits using summary statistics from biobank-scale genome-wide association studies (GWAS). However, empirical evidence has shown a common bias phenomenon that highly significant cross-trait PRS can only account for a very small amount of genetic variance (<i>R</i><sup>2</sup> can be <1%) in independent testing GWAS. The aim of this paper is to investigate and address the bias phenomenon of cross-trait PRS in numerous GWAS applications. We show that the estimated genetic correlation can be asymptotically biased toward zero. A consistent cross-trait PRS estimator is then proposed to correct such asymptotic bias. In addition, we investigate whether or not SNP screening by GWAS <i>p</i>-values can lead to improved estimation and show the effect of overlapping samples among GWAS. We analyze GWAS summary statistics of reaction time and brain structural magnetic resonance imaging-based features measured in the Pediatric Imaging, Neurocognition, and Genetics study. We find that the raw cross-trait PRS estimators heavily underestimate the genetic similarity between cognitive function and human brain structures (mean R2=1.32%), whereas the bias-corrected estimators uncover the moderate degree of genetic overlap between these closely related heritable traits (mean R2=22.42%). Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.
研究动机与目标
- 为解决一个持续存在的经验现象:即在存在强烈遗传重叠的情况下,高度显著的交叉特征PRS仍解释不到1%的遗传方差。
- 从理论上证明,在高度多基因或全基因组结构下,基于PRS的遗传相关性估计量会渐近地偏向零。
- 构建一种一致的PRS估计量,以消除该渐近偏差,且独立于未知的因果SNP数量。
- 提出一种仅依赖GWAS汇总统计量、无需个体水平数据的新遗传相关性估计量。
- 研究通过p值筛选SNP以及重叠样本对估计精度和偏差的影响。
提出的方法
- 在包含全部p个SNP的多基因模型下对PRS估计进行理论分析,表明渐近偏差仅取决于(n, p, h²),与因果SNP数量m无关。
- 推导出一种一致的PRS估计量,通过调整高维设置下非因果SNP的影响,从而消除渐近偏差。
- 基于两组GWAS汇总统计量,开发一种新型遗传相关性估计量,利用PRS预测中的交叉特征协方差和方差分量。
- 在维度和样本量同时增加的高维渐近理论框架下进行分析,对mαβ、mα、mβ和p随n增长的条件作出假设。
- 应用随机矩阵理论和集中不等式,推导检验统计量和相关性估计量的渐近行为。
- 通过数值模拟和对英国生物银行大脑白质束及神经精神疾病数据的分析,进行有限样本验证。
实验结果
研究问题
- RQ1为何在高度多基因特征中,即使关联高度显著,交叉特征PRS方法仍系统性地低估遗传相关性?
- RQ2在多基因或全基因组模型下,基于PRS的遗传相关性估计中渐近偏差的理论根源是什么?
- RQ3能否构建一种一致的PRS估计量,使其在未知因果SNP数量的情况下,完全消除渐近偏差?
- RQ4通过GWAS p值进行SNP筛选是否能提高高维设置下遗传相关性估计的精度?
- RQ5不同GWAS数据集中样本重叠如何影响遗传相关性估计的偏差与方差?
主要发现
- 当特征为高度多基因或全基因组结构时,标准交叉特征PRS估计的遗传相关性会渐近地偏向零,即使样本量很大。
- PRS估计量的渐近偏差与因果SNP数量m无关,仅取决于三重组合(n, p, h²),其中n为样本量,p为SNP数量,h²为SNP遗传力。
- 所提出的稳定PRS估计量成功消除了该渐近偏差,使在高维设置下准确估计遗传相关性成为可能。
- 基于仅GWAS汇总统计量的新遗传相关性估计量在相同渐近框架下实现了稳定估计。
- 数值实验和英国生物银行数据分析证实,偏差校正后的PRS估计量显著提高了遗传重叠预测的R²,解决了“遗传重叠缺失”悖论。
- 通过p值筛选SNP并不能提高估计精度,反而可能加剧偏差;而重叠样本会引入额外偏差,必须在分析中予以校正。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。