Skip to main content
QUICK REVIEW

[论文解读] Robust Inference Using Inverse Probability Weighting

Xinwei Ma, Jingshen Wang|arXiv (Cornell University)|Oct 26, 2018
Statistical Methods and Inference参考文献 22被引用 7
一句话总结

本文提出了一种针对逆概率加权(IPW)估计量的稳健推断程序,该程序可适应由小概率权重和修剪阈值引起的非高斯渐近分布。通过结合子采样与局部多项式偏差校正,并采用数据驱动的修剪阈值选择方法,该方法在各种尾部行为和修剪选择下均能确保有效推断,即使标准高斯近似失效亦成立。

ABSTRACT

Inverse Probability Weighting (IPW) is widely used in empirical work in economics and other disciplines. As Gaussian approximations perform poorly in the presence of "small denominators," trimming is routinely employed as a regularization strategy. However, ad hoc trimming of the observations renders usual inference procedures invalid for the target estimand, even in large samples. In this paper, we first show that the IPW estimator can have different (Gaussian or non-Gaussian) asymptotic distributions, depending on how "close to zero" the probability weights are and on how large the trimming threshold is. As a remedy, we propose an inference procedure that is robust not only to small probability weights entering the IPW estimator but also to a wide range of trimming threshold choices, by adapting to these different asymptotic distributions. This robustness is achieved by employing resampling techniques and by correcting a non-negligible trimming bias. We also propose an easy-to-implement method for choosing the trimming threshold by minimizing an empirical analogue of the asymptotic mean squared error. In addition, we show that our inference procedure remains valid with the use of a data-driven trimming threshold. We illustrate our method by revisiting a dataset from the National Supported Work program.

研究动机与目标

  • 解决当概率权重接近零时,标准高斯近似在IPW推断中失效的问题。
  • 克服因人为修剪导致的传统推断程序无效的问题,该问题会引入显著偏差并改变渐近分布。
  • 开发一种双向稳健推断方法,可适应重尾概率权重和不同的修剪阈值。
  • 提出一种基于最小化经验渐近均方误差的自适应修剪阈值选择规则。
  • 确保即使修剪阈值基于数据驱动标准内生选择,推断仍然有效。

提出的方法

  • 使用子采样构建临界值,使其能够适应在不同尾部分布和修剪阈值下IPW估计量的未知渐近分布。
  • 实施一种新颖的基于局部多项式的偏差校正方法,以校正与权重低于阈值概率成比例的显著修剪偏差。
  • 通过最小化渐近均方误差的经验近似形式,提出一种数据驱动的修剪阈值选择规则。
  • 将IPW估计量与修剪阈值 $ b_n $ 联合定义,排除估计倾向得分 $ \hat{e}(X_i) < b_n $ 的观测值,以对小权重进行正则化。
  • 推导修剪后IPW估计量的渐近分布,表明其可能为非高斯分布,且在权重在零附近呈现重尾时收敛速度慢于 $ \sqrt{n} $。
  • 结合重采样与偏差校正,构建在权重分布和修剪水平不同配置下均保持有效的稳健置信区间。

实验结果

研究问题

  • RQ1当概率权重接近零并应用修剪时,IPW估计量的渐近分布如何变化?
  • RQ2修剪对IPW估计量的渐近偏差和收敛速度有何影响?
  • RQ3能否开发一种统一的推断程序,使其在倾向得分的不同尾部分布和不同修剪阈值下均保持有效?
  • RQ4如何选择修剪阈值,以最小化渐近均方误差,同时保持推断的有效性?
  • RQ5在具有重尾权重的小样本中,所提出的方法相较于基于标准高斯分布的推断在多大程度上表现更优?

主要发现

  • 当概率权重在零附近呈现重尾时,IPW估计量可能具有非高斯渐近分布,且收敛速度慢于 $ \sqrt{n} $。
  • 修剪会引入显著的渐近偏差,该偏差取决于权重低于阈值的概率,且即使在大样本中也不会消失。
  • 渐近分布依赖于三个参数:尾指数 $ \gamma_0 $ 和形状参数 $ \alpha_+ $ 与 $ \alpha_- $,可能导致非对称且非正态的抽样分布。
  • 所提出的结合子采样与偏差校正的推断程序,可产生对底层渐近分布或修剪阈值选择不敏感的稳健置信区间。
  • 基于最小化经验渐近均方误差的自适应修剪阈值选择规则,可实现稳定且高效的推断,实证应用中仅修剪了五个观测值。
  • 在国家支持工作数据集中,稳健置信区间呈现非对称性,且比传统高斯近似区间更可靠,后者过于狭窄且对修剪选择敏感。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。