Skip to main content
QUICK REVIEW

[论文解读] Estimation and Inference for High Dimensional Generalized Linear Models: A Splitting and Smoothing Approach

Zhe Fei, Yi Li|arXiv (Cornell University)|Mar 11, 2019
Statistical Methods and Inference参考文献 24被引用 12
一句话总结

该论文提出了一种计算高效的‘分割与平滑’方法,用于高维广义线性模型(GLMs),通过数据分割实现有效推断:在一部分数据上进行变量选择,在另一部分上进行低维估计,随后对多次分割的结果进行平均。该方法在无需估计高维精度矩阵的情况下,实现了渐近正态性与正确的置信区间覆盖。

ABSTRACT

The focus of modern biomedical studies has gradually shifted to explanation and estimation of joint effects of high dimensional predictors on disease risks. Quantifying uncertainty in these estimates may provide valuable insight into prevention strategies or treatment decisions for both patients and physicians. High dimensional inference, including confidence intervals and hypothesis testing, has sparked much interest. While much work has been done in the linear regression setting, there is lack of literature on inference for high dimensional generalized linear models. We propose a novel and computationally feasible method, which accommodates a variety of outcome types, including normal, binomial, and Poisson data. We use a "splitting and smoothing" approach, which splits samples into two parts, performs variable selection using one part and conducts partial regression with the other part. Averaging the estimates over multiple random splits, we obtain the smoothed estimates, which are numerically stable. We show that the estimates are consistent, asymptotically normal, and construct confidence intervals with proper coverage probabilities for all predictors. We examine the finite sample performance of our method by comparing it with the existing methods and applying it to analyze a lung cancer cohort study.

研究动机与目标

  • 解决高维广义线性模型(GLMs)缺乏可靠推断方法的问题,特别是针对二值或计数等复杂结果。
  • 克服现有方法在计算和理论上面临的挑战,这些方法需要对大型精度矩阵求逆,或依赖严格的调参选择。
  • 开发一种方法,确保在高维稀疏条件下,置信区间和假设检验具有正确的覆盖概率。
  • 通过在多次随机数据分割上取平均,提升估计的稳定性和效率,降低选择与分割带来的变异性。
  • 使方法可实际应用于真实世界生物医学数据,如肺癌研究中的SNP与基因互作分析,在稀疏条件下实现稳健推断。

提出的方法

  • 将数据集随机分为两部分:一部分用于变量选择,另一部分用于估计。
  • 使用第一部分进行变量选择(例如通过LASSO),以识别重要预测变量,实现维度降低。
  • 在第二部分上拟合低维GLM,通过依次添加每个预测变量(无论是否被选中)来估计其系数。
  • 重复分割与估计过程B次,然后对得到的系数估计值取平均,获得平滑且稳定的估计量。
  • 应用微小自助法(infinitesimal jackknife)推导出能考虑分割与选择变异性影响的稳健方差估计量。
  • 利用平滑后的估计量及其方差-协方差矩阵构造置信区间并进行假设检验,确保在弱条件下具有正确的覆盖性。

实验结果

研究问题

  • RQ1能否开发一种计算上可行的方法,用于高维GLMs的推断,且无需估计高维精度矩阵?
  • RQ2在稀疏条件下,分割与平滑方法是否能产生渐近正态的估计量并具有有效的置信区间?
  • RQ3与现有的去稀疏化LASSO和后选择推断方法相比,该方法在有限样本下的表现如何?
  • RQ4该方法能否在存在复杂交互作用(如遗传流行病学中的SNP-吸烟交互作用)时,实现可靠的推断?
  • RQ5当稀疏性假设仅部分满足或轻微违反时,该方法的稳健性如何?

主要发现

  • 所提出的SSGLM方法即使在高维设置下,也能实现渐近正态性与正确的置信区间覆盖概率。
  • 在肺癌队列研究中,使用SSGLM的Model 2解释了疾病状态25%更多的变异(R² = 0.1168),优于基线模型(R² = 0.0938),也优于去稀疏化LASSO方法(R² = 0.1018)。
  • 基于SSGLM的模型AUC为0.69,高于去稀疏化LASSO模型的0.668,表明其具有更好的判别性能。
  • 对于SNP rs3117582,该方法未发现其与肺癌存在显著关联(p值 = 0.97),解决了先前研究中相互矛盾的报告。
  • 该方法通过支持数据分割与低维GLM拟合的并行化,展现出优越的计算效率。
  • 微小自助法方差估计量在无需参数假设的情况下提供了准确的推断,其稳定性和鲁棒性类似于集成方法(bagging)的特性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。