[论文解读] Simultaneous inference for generalized linear models with unmeasured confounders
本文提出了一种统一框架,用于在存在未测量混杂因素的多元广义线性模型中进行联合推断,通过正交结构和投影偏差校正,利用Lasso型优化恢复潜在系数并估计主要效应。该方法确保了渐近有效的第一类错误控制和有效的错误发现率(FDR)控制,在高维、非正态响应设置下,其统计功效和稳健性优于现有方法。
Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may be substantially biased. This paper investigates the large-scale hypothesis testing problem for multivariate generalized linear models in the presence of confounding effects. Under arbitrary confounding mechanisms, we propose a unified statistical estimation and inference framework that harnesses orthogonal structures and integrates linear projections into three key stages. It begins by disentangling marginal and uncorrelated confounding effects to recover the latent coefficients. Subsequently, latent factors and primary effects are jointly estimated through lasso-type optimization. Finally, we incorporate projected and weighted bias-correction steps for hypothesis testing. Theoretically, we establish the identification conditions of various effects and non-asymptotic error bounds. We show effective Type-I error control of asymptotic $z$-tests as sample and response sizes approach infinity. Numerical experiments demonstrate that the proposed method controls the false discovery rate by the Benjamini-Hochberg procedure and is more powerful than alternative methods. By comparing single-cell RNA-seq counts from two groups of samples, we demonstrate the suitability of adjusting confounding effects when significant covariates are absent from the model.
研究动机与目标
- 解决在基因组学中大规模假设检验的多元广义线性模型中存在未测量混杂因素的挑战。
- 开发一种统一的统计框架,识别并调整混杂效应,而无需了解观测协变量与潜在因子之间的因果关系。
- 在任意混杂机制和非正态响应分布下,确保有效的统计推断——特别是第一类错误控制和错误发现率(FDR)控制。
- 将现有方法从线性模型和单变量结果设置扩展到处理现代组学研究中常见的高维、多变量和非线性数据。
提出的方法
- 通过正交分解分离边际效应与不相关混杂效应,实现对潜在系数的一致估计。
- 采用Lasso型优化框架联合估计潜在因子和主要效应,对系数矩阵和因子载荷均施加正则化。
- 集成投影和加权偏差校正步骤,以校正主要效应估计中的偏差,从而实现有效的渐近z检验。
- 利用线性投影和非渐近分析,建立在任意混杂机制下的可识别性条件和误差界。
- 在推断阶段使用样本分割,以确保估计与检验之间的独立性,尽管即使不使用样本分割,该方法仍保持稳健。
- 该框架适用于具有多变量响应的广义线性模型,可处理非正态分布,如单细胞RNA-seq数据中的分布。
实验结果
研究问题
- RQ1当存在未测量混杂因素且未知其与观测协变量之间的因果关系时,我们能否在多元广义线性模型中实现有效的联合推断?
- RQ2在非正态、稀疏或过度分散的响应下,我们如何在高维设置中一致估计主要效应和潜在混杂因素?
- RQ3所提出的方法在任意混杂机制下是否能保持渐近的第一类错误控制和有效的错误发现率(FDR)控制?
- RQ4在真实世界基因组数据中,该方法在统计功效和稳健性方面相较于现有方法的优越程度如何?
主要发现
- 在Benjamini-Hochberg程序下,所提出的方法有效控制了错误发现率(FDR),模拟中错误发现比例(FDP)的中位数为0.191–0.219。
- 该方法具有很高的统计功效,当不使用样本分割时,中位功效达到0.987;当保留80%的数据用于推断时,中位功效为0.963。
- 第一类错误得到良好控制,所有模拟设置下的中位率在0.050–0.051之间,表明z检验具有渐近有效性。
- 建立了对估计系数矩阵的非渐近误差界,满足$\rVert\bm{B} - \bm{B}^*\rVert_{\text{F}} \lesssim 1/\sqrt{n \wedge p}$,表明在高维情形下估计具有一致性。
- 该方法对样本分割比例具有鲁棒性,即使仅保留20%的数据用于推断,性能下降也微乎其微。
- 在系统性红斑狼疮患者单细胞RNA-seq数据中,即使关键协变量未被观测到,该方法仍能成功调整混杂效应,展示了其在真实世界基因组学应用中的实际效用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。