[论文解读] Honest Confidence Regions for Logistic Regression with a Large Number of Controls
该论文提出了一种稳健的方法,用于在控制变量数量超过样本量时,估计并构建逻辑回归中感兴趣系数的诚实置信区域。通过利用稀疏性和工具变量技术,该方法在弱正则性条件下实现了根n估计和一致有效性,且无需依赖一致的模型选择。
This paper considers generalized linear models in the presence of many controls. We lay out a general methodology to estimate an effect of interest based on the construction of an instrument that immunize against model selection mistakes and apply it to the case of logistic binary choice model. More specifically we propose new methods for estimating and constructing confidence regions for a regression parameter of primary interest $\alpha_0$, a parameter in front of the regressor of interest, such as the treatment variable or a policy variable. These methods allow to estimate $\alpha_0$ at the root-$n$ rate when the total number $p$ of other regressors, called controls, potentially exceed the sample size $n$ using sparsity assumptions. The sparsity assumption means that there is a subset of $s<n$ controls which suffices to accurately approximate the nuisance part of the regression function. Importantly, the estimators and these resulting confidence regions are valid uniformly over $s$-sparse models satisfying $s^2\log^2 p = o(n)$ and other technical conditions. These procedures do not rely on traditional consistent model selection arguments for their validity. In fact, they are robust with respect to moderate model selection mistakes in variable selection. Under suitable conditions, the estimators are semi-parametrically efficient in the sense of attaining the semi-parametric efficiency bounds for the class of models in this paper.
研究动机与目标
- 解决当控制变量数量超过样本量时,在逻辑回归中估计感兴趣参数的挑战。
- 开发即使在控制变量模型选择不完美或不一致时仍保持有效的推断程序。
- 确保在一大类高维稀疏模型中,估计和置信区域具有一致有效性。
- 在不依赖一致模型选择或强参数假设的前提下,实现半参数效率。
提出的方法
- 构建一个与控制变量模型选择误差不相关的工具变量。
- 利用稀疏性假设:仅需一小部分控制变量(s < n)即可准确近似干扰回归函数。
- 通过两阶段程序估计感兴趣参数,将主效应估计与干扰函数估计分离。
- 应用去偏技术以校正因高维控制选择引入的估计偏差。
- 通过确保在s-稀疏模型上的一致有效性,构建对中等程度模型选择错误具有鲁棒性的置信区域。
- 在条件 s² log²p = o(n) 下利用渐近理论,以确保估计量的根n收敛性和渐近正态性。
实验结果
研究问题
- RQ1当控制变量数量超过样本量时,能否为逻辑回归中感兴趣参数构建诚实置信区域?
- RQ2当控制变量模型选择不一致或不完美时,如何确保推断的有效性?
- RQ3在什么条件下可实现高维逻辑模型中感兴趣参数的根n估计?
- RQ4所提出的方法能否在不依赖一致模型选择的前提下实现半参数效率?
- RQ5该方法在满足 s² log²p = o(n) 的不同稀疏模型中表现如何?
主要发现
- 所提出的估计量即使在控制变量数量 p 超过样本量 n 时,也能实现感兴趣参数的根n收敛。
- 在条件 s² log²p = o(n) 下,使用该方法构建的置信区域在s-稀疏模型上具有一致有效性。
- 该方法在中等程度的模型选择错误下仍保持有效,且无需控制变量的一致选择。
- 在适当的正则性条件下,估计量达到半参数效率下界,表明其具有最优估计性能。
- 该方法不依赖于一致模型选择来保证有效性,因此对高维变量选择中的常见陷阱具有鲁棒性。
- 该方法可广泛应用于具有大量控制变量的广义线性模型,特别聚焦于逻辑回归二元选择模型。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。