Skip to main content
QUICK REVIEW

[论文解读] Estimation, Confidence Intervals, and Large-Scale Hypotheses Testing for High-Dimensional Mixed Linear Regression

Linjun Zhang, Rong Ma|ArXiv.org|Nov 6, 2020
Statistical Methods and Inference参考文献 36被引用 7
一句话总结

本文提出了一种高维混合线性回归框架,其中混合比例和协方差结构未知,采用高维EM算法进行估计,利用去偏估计量实现渐近正态性,并设计大规模多重检验程序以控制FDR。该方法可在高维异质数据设置下实现有效的置信区间和假设检验。

ABSTRACT

This paper studies the high-dimensional mixed linear regression (MLR) where the output variable comes from one of the two linear regression models with an unknown mixing proportion and an unknown covariance structure of the random covariates. Building upon a high-dimensional EM algorithm, we propose an iterative procedure for estimating the two regression vectors and establish their rates of convergence. Based on the iterative estimators, we further construct debiased estimators and establish their asymptotic normality. For individual coordinates, confidence intervals centered at the debiased estimators are constructed. Furthermore, a large-scale multiple testing procedure is proposed for testing the regression coefficients and is shown to control the false discovery rate (FDR) asymptotically. Simulation studies are carried out to examine the numerical performance of the proposed methods and their superiority over existing methods. The proposed methods are further illustrated through an analysis of a dataset of multiplex image cytometry, which investigates the interaction networks among the cellular phenotypes that include the expression levels of 20 epitopes or combinations of markers.

研究动机与目标

  • 解决在混合比例和协变量协方差结构未知的高维混合线性回归(MLR)中缺乏理论与计算方法的问题。
  • 为p >> n的高维MLR开发一种计算高效且理论基础坚实的估计方法。
  • 为单个回归系数及其差异构建渐近有效的置信区间。
  • 设计一种大规模多重检验程序,以在高维设置下控制错误发现率(FDR)和错误发现比例(FDP)。
  • 实现在包含异质亚群的复杂生物数据集(如多重图像细胞术数据)中的推断。

提出的方法

  • 提出一种迭代高维EM算法,用于在未知混合比例ω*和协方差矩阵Σ的条件下估计两个回归向量β₁*和β₂*。
  • 推导出迭代估计量,并在稀疏性和高维条件下建立其收敛速率。
  • 通过校正迭代估计量中的偏差,构造去偏估计量,以实现渐近正态性。
  • 利用去偏估计量的渐近正态性,以去偏估计为中心构建单个回归系数的置信区间。
  • 基于去偏检验统计量,开发一种大规模多重检验程序,以渐近方式控制FDR和FDP。
  • 采用基于标准正态分布分位数的阈值规则确定拒绝域,确保在高维渐近条件下实现FDR控制。

实验结果

研究问题

  • RQ1当混合比例ω*和设计协方差Σ均未知时,如何高效估计高维混合线性回归中两个回归向量β₁*和β₂*?
  • RQ2如何为β₁*和β₂*的各个分量构建渐近有效的置信区间?
  • RQ3如何设计一种大规模多重检验程序,以控制对p个预测变量进行H₀j: β₁j* = β₂j* = 0检验时的错误发现率(FDR)?
  • RQ4在高维渐近条件下,迭代估计量的理论收敛速率是什么?
  • RQ5所提出的方法在有限样本下表现如何,尤其是在真实生物数据设置下与现有方法相比的表现如何?

主要发现

  • 在稀疏性条件s = o(n^{1/2}/(log^{3/2}p log n))下,β₁*和β₂*的迭代估计量达到最优收敛速率。
  • 去偏估计量具有渐近正态性,从而可构建单个回归系数的有效置信区间。
  • 所提出的大型规模多重检验程序在高维渐近条件下能渐近控制错误发现率(FDR)和错误发现比例(FDP)。
  • 在模拟研究中,该方法相较于现有方法展现出更优的数值性能,尤其在高维和异质性设置下。
  • 该方法成功应用于多重图像细胞术数据集,揭示了20个细胞标记物之间具有生物学意义的相互作用网络。
  • 在最小假设条件下建立了理论保证,包括未知混合比例ω* ∈ (1/2, 1)和未知协方差结构Σ。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。