Skip to main content
QUICK REVIEW

[论文解读] Derandomizing Knockoffs

Zhimei Ren, Yuting Wei|arXiv (Cornell University)|Dec 4, 2020
Genetic Associations and Epidemiology参考文献 61被引用 6
一句话总结

本文提出了一种 Model-X knockoffs 的去随机化框架,通过聚合多次运行的结果,实现稳定且一致的变量选择,同时保持严格的错误控制。通过结合 knockoffs 与稳定性选择原则,该方法实现了家庭错误率(PFER)和 k 重家庭错误率(k-FWER)控制,在高维变量选择中显著提升了统计功效和 I 类错误控制能力,尤其适用于多阶段全基因组关联研究(GWAS)。

ABSTRACT

Model-X knockoffs is a general procedure that can leverage any feature importance measure to produce a variable selection algorithm, which discovers true effects while rigorously controlling the number or fraction of false positives. Model-X knockoffs is a randomized procedure which relies on the one-time construction of synthetic (random) variables. This paper introduces a derandomization method by aggregating the selection results across multiple runs of the knockoffs algorithm. The derandomization step is designed to be flexible and can be adapted to any variable selection base procedure to yield stable decisions without compromising statistical power. When applied to the base procedure of Janson et al. (2016), we prove that derandomized knockoffs controls both the per family error rate (PFER) and the k family-wise error rate (k-FWER). Further, we carry out extensive numerical studies demonstrating tight type-I error control and markedly enhanced power when compared with alternative variable selection algorithms. Finally, we apply our approach to multi-stage genome-wide association studies of prostate cancer and report locations on the genome that are significantly associated with the disease. When cross-referenced with other studies, we find that the reported associations have been replicated.

研究动机与目标

  • 为解决随机 knockoffs 的不稳定性问题,其在不同运行中因随机生成 knockoff 而导致变量选择结果不一致。
  • 开发一种通用的去随机化框架,在不损害统计功效或错误率保证的前提下提升稳定性。
  • 在确证性研究中实现可靠的变量选择,特别是在对严格错误控制和可重复性要求极高的多阶段 GWAS 中。
  • 提供一种对聚合发现集实现 I 类错误控制的方法,克服以往 knockoffs 应用中基于选择频率报告的局限性。
  • 通过真实基因组数据验证该方法的有效性,实现跨独立研究的生物学可重复发现。

提出的方法

  • 使用独立的随机种子多次运行 Model-X knockoffs 算法,生成多个选择集合。
  • 采用受稳定性选择启发的准则聚合选择结果,选择在多次运行中一致识别出的变量。
  • 对选择频率设定阈值,形成最终的发现集合,确保稳定性并减少假阳性。
  • 利用 Janson 等人(2016)的理论结果,证明聚合过程可同时控制 PFER 和 k-FWER。
  • 设计两阶段 GWAS 工作流程:首先,使用去随机化的 knockoffs 识别候选 SNP;其次,在后续研究中验证发现结果。
  • 在 SNP 聚类中使用 2% 分辨率,将邻近变异分组,并报告每个簇的主导 SNP,以便与以往文献进行比较。

实验结果

研究问题

  • RQ1能否构建一种去随机化的 knockoffs 方法,在保持严格错误控制的同时提升选择稳定性?
  • RQ2聚合多个 knockoff 运行是否相比标准 knockoffs 提升了更高的统计功效?
  • RQ3当基础方法具备 PFER 控制能力时,去随机化 knockoffs 方法是否能同时控制 PFER 和 k-FWER?
  • RQ4去随机化 knockoffs 的发现是否在独立的 GWAS 研究中具有可重复性并得到生物学验证?
  • RQ5该方法是否可适配于多阶段 GWAS 工作流程,以支持具有强错误率保证的确认性分析?

主要发现

  • 去随机化 knockoffs 即使在基础方法仅控制 PFER 的情况下,也能对 PFER 和 k-FWER 实现紧密控制。
  • 在数值研究中,该方法相比原始 knockoffs 及其他变量选择算法,表现出显著增强的统计功效。
  • 在 2% 分辨率和 FWER 水平 0.1 下发现的所有 SNP,均已在文献中被先前报道,或在独立 GWAS 研究中得到验证。
  • 当 FWER 阈值放宽以控制 3-FWER 在 0.1 水平时,额外发现了七个 SNP,全部得到先前文献或荟萃分析的支持。
  • 每个簇的主导 SNP(如 rs12621278、rs1016343、rs7501939)均在外部研究中被发现具有生物学可重复性,验证了该方法的可靠性。
  • 该方法支持科学合理的两阶段 GWAS 工作流程:候选 SNP 在严格错误控制下被发现,并在后续研究中得到确认。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。