[论文解读] High Dimensional M-Estimation with Missing Outcomes: A Semi-Parametric Framework
本文提出了一种半参数、高维的M-估计框架,适用于存在缺失结果和高维协变量的情形,引入了一种L1-正则化去偏双重稳健(DDR)估计器。当倾向得分模型和结果回归模型均正确时,该方法可实现最优的$L_2$误差率$\sqrt{s(\log d)/n}$;当仅其中一个模型正确时,仍保持一致性和双重稳健性,从而在模型误设情况下也能实现有效的推断。
We consider high dimensional $M$-estimation in settings where the response $Y$ is possibly missing at random and the covariates $\mathbf{X} \in \mathbb{R}^p$ can be high dimensional compared to the sample size $n$. The parameter of interest $\boldsymbolθ_0 \in \mathbb{R}^d$ is defined as the minimizer of the risk of a convex loss, under a fully non-parametric model, and $\boldsymbolθ_0$ itself is high dimensional which is a key distinction from existing works. Standard high dimensional regression and series estimation with possibly misspecified models and missing $Y$ are included as special cases, as well as their counterparts in causal inference using 'potential outcomes'. Assuming $\boldsymbolθ_0$ is $s$-sparse ($s \ll n$), we propose an $L_1$-regularized debiased and doubly robust (DDR) estimator of $\boldsymbolθ_0$ based on a high dimensional adaptation of the traditional double robust (DR) estimator's construction. Under mild tail assumptions and arbitrarily chosen (working) models for the propensity score (PS) and the outcome regression (OR) estimators, satisfying only some high-level conditions, we establish finite sample performance bounds for the DDR estimator showing its (optimal) $L_2$ error rate to be $\sqrt{s (\log d)/ n}$ when both models are correct, and its consistency and DR properties when only one of them is correct. Further, when both the models are correct, we propose a desparsified version of our DDR estimator that satisfies an asymptotic linear expansion and facilitates inference on low dimensional components of $\boldsymbolθ_0$. Finally, we discuss various of choices of high dimensional parametric/semi-parametric working models for the PS and OR estimators. All results are validated via detailed simulations.
研究动机与目标
- 解决在参数本身为高维时的高维M-估计问题,尤其在存在缺失结果的情形下。
- 开发一种双重稳健、去偏且正则化的估计器,当倾向得分或结果回归模型其中之一正确时仍能保持一致性。
- 在较弱的尾部和高层级建模假设下,建立所提估计器的有限样本性能界。
- 通过DDR估计器的去稀疏化版本,实现对高维参数$\bm{\theta}_0$的低维分量的有效推断。
- 为干扰函数提供灵活的高维参数或半参数工作模型选择框架。
提出的方法
- 通过将经典双重稳健估计器适配到高维稀疏设定,提出一种L1-正则化去偏双重稳健(DDR)估计器。
- 采用高维双重稳健构造的适应版本,即在倾向得分或结果回归模型其中之一正确时,估计器具有双重稳健性。
- 在较弱的尾部假设和对倾向得分与结果回归工作模型的高层级条件下,建立有限样本$L_2$误差界。
- 提出DDR估计器的去稀疏化版本,其满足渐近线性展开,从而实现渐近正态性,并可对$\bm{\theta}_0$的低维分量进行推断。
- 依赖高维级数估计和正则化技术,以处理本身稀疏($s \ll n$个非零条目)的高维参数$\bm{\theta}_0$的高维特性。
- 采用半参数框架,允许对潜在风险函数进行灵活的非参数建模,同时保持有限样本性能保证。
实验结果
研究问题
- RQ1能否构建一种高维M-估计器,在缺失结果和高维协变量下仍保持最优的$L_2$误差率?
- RQ2当仅倾向得分或结果回归模型其中之一正确时,所提出的估计器是否仍保持双重稳健性?
- RQ3在模型误设和缺失数据条件下,能否对高维参数$\bm{\theta}_0$的低维分量进行有效推断?
- RQ4在较弱的尾部假设和建模假设下,所提DDR估计器的有限样本性能界是什么?
- RQ5对干扰函数采用不同的高维参数或半参数模型,如何影响估计器的性能?
主要发现
- 当倾向得分和结果回归模型均正确时,所提出的DDR估计器实现了最优的$L_2$误差率$\sqrt{s(\log d)/n}$。
- 当仅两个模型其中之一正确时,估计器仍保持一致性和双重稳健性,确保在模型误设情况下的有效推断。
- DDR估计器的去稀疏化版本满足渐近线性展开,从而实现渐近正态性,并可对$\bm{\theta}_0$的低维分量进行有效推断。
- 在较弱的尾部假设和对工作模型的高层级条件下,建立了有限样本性能界,确保对模型误设的鲁棒性。
- 通过详细模拟验证了该框架,结果表明估计器在各种高维和缺失数据场景下均表现出鲁棒性和高效性。
- 该方法适用于标准高维回归、级数估计以及潜在结果的因果推断,扩展了现有框架至高维参数情形。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。