[论文解读] Narrowing the gap on heritability of common disease by direct estimation in case-control GWAS
本文提出遗传相关性回归(GCR),一种直接估计病例对照全基因组关联研究(GWAS)遗传力的方法,通过将表型相似性对基因型相关性进行回归,同时校正选择性偏差。与现有线性混合模型(LMM)方法相比,GCR可得到无偏且更高的遗传力估计值——例如克罗恩病的遗传力为34%,而LMM方法仅为22%——并引入一种启发式校正方法,将多发性硬化症的遗传力估计值从30%提高至52.6%。
One of the major developments in recent years in the search for missing heritability of human phenotypes is the adoption of linear mixed-effects models (LMMs) to estimate heritability due to genetic variants which are not significantly associated with the phenotype. A variant of the LMM approach has been adapted to case-control studies and applied to many major diseases by Lee et al. (2011), successfully accounting for a considerable portion of the missing heritability. For example, for Crohn's disease their estimated heritability was 22% compared to 50-60% from family studies. In this letter we propose to estimate heritability of disease directly by regression of phenotype similarities on genotype correlations, corrected to account for ascertainment. We refer to this method as genetic correlation regression (GCR). Using GCR we estimate the heritability of Crohn's disease at 34% using the same data. We demonstrate through extensive simulation that our method yields unbiased heritability estimates, which are consistently higher than LMM estimates. Moreover, we develop a heuristic correction to LMM estimates, which can be applied to published LMM results. Applying our heuristic correction increases the estimated heritability of multiple sclerosis from 30% to 52.6%.
研究动机与目标
- 弥合常见疾病中基于家系的遗传力估计与GWAS估计的遗传力之间的长期差距。
- 克服线性混合模型(LMM)在选择性病例对照研究中的局限性,因为病例过度代表会违反正态性和独立性假设。
- 开发一种直接、无偏的方法,利用选择样本中的基因型相关性和表型相似性来估计遗传力。
- 提供一种实用的启发式校正方法,用于校正先前发表的基于LMM的遗传力估计值中的选择性偏差。
- 通过改进疾病GWAS中的遗传力估计,提高多基因风险预测和遗传结构推断的准确性。
提出的方法
- 提出遗传相关性回归(GCR),通过将成对表型相似性对成对基因型相关性进行回归,直接估计遗传力。
- 通过使用推导出的解析表达式来校正选择性抽样下真实遗传方差的期望值,从而校正选择性偏差。
- 利用模拟验证GCR在不同疾病患病率和遗传方差水平下均能产生无偏的遗传力估计。
- 基于选择性下真实遗传方差与LMM估计方差的比值,推导出一种启发式校正因子,使先前基于LMM的估计结果可被校正。
- 通过反向推导估计偏差,将该校正方法应用于真实数据(如多发性硬化症),以恢复校正后的遗传力估计值。
- 使用GCTA软件计算基于LMM的遗传力估计值,并仅利用观测估计值、疾病患病率(P)和病例比例(K)对已发表的LMM结果应用该校正。
实验结果
研究问题
- RQ1基于回归的直接方法是否能在选择性病例对照GWAS中比现有LMM方法更准确地估计遗传力?
- RQ2选择性偏差如何系统性地扭曲病例对照研究中基于LMM的遗传力估计?
- RQ3能否推导出一种启发式校正方法,用于校正先前发表的基于LMM的遗传力估计中的选择性偏差?
- RQ4在不同病例比例和疾病患病率下,LMM估计值的偏差程度如何?
- RQ5GCR是否在一系列模拟的遗传结构和选择性水平下均能产生无偏的遗传力估计?
主要发现
- GCR在所有模拟情景下均产生无偏的遗传力估计,包括存在强烈选择性偏差的情况。
- GCR估计值始终高于LMM估计值——例如,克罗恩病的遗传力为34%,而使用相同数据的LMM方法仅为22%。
- 当应用于已发表结果时,启发式校正方法将多发性硬化症的遗传力估计值从30%(LMM)提高至52.6%。
- 当病例比例与疾病患病率之比(K/P)减小时,LMM估计值的偏差加剧,表明在罕见病研究中偏差更大。
- 校正因子呈非线性关系,其推导基于模拟得到的观测LMM估计值与真实遗传力之间的关系。
- 校正后估计值的置信区间通过将相同的校正程序应用于原始95%置信区间的上下限来获得。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。