[论文解读] Semiparametric Estimation with Data Missing Not at Random Using an Instrumental Variable
本文提出了一种半参数框架,用于在缺失非随机(MNAR)数据下,使用工具变量(IV)对总体均值进行非参数识别与估计。该研究开发了一种新颖的双重稳健估计量,结合了逆概率加权与结果回归,确保若倾向得分或结果模型任一正确设定,估计结果均具一致性,其应用实例为博茨瓦纳的HIV血清流行率研究,其中访谈员特征被用作IV。
Missing data occur frequently in empirical studies in health and social sciences, often compromising our ability to make accurate inferences. An outcome is said to be missing not at random (MNAR) if, conditional on the observed variables, the missing data mechanism still depends on the unobserved outcome. In such settings, identification is generally not possible without imposing additional assumptions. Identification is sometimes possible, however, if an instrumental variable (IV) is observed for all subjects which satisfies the exclusion restriction that the IV affects the missingness process without directly influencing the outcome. In this paper, we provide necessary and sufficient conditions for nonparametric identification of the full data distribution under MNAR with the aid of an IV. In addition, we give sufficient identification conditions that are more straightforward to verify in practice. For inference, we focus on estimation of a population outcome mean, for which we develop a suite of semiparametric estimators that extend methods previously developed for data missing at random. Specifically, we propose inverse probability weighted estimation, outcome regression-based estimation and doubly robust estimation of the mean of an outcome subject to MNAR. For illustration, the methods are used to account for selection bias induced by HIV testing refusal in the evaluation of HIV seroprevalence in Mochudi, Botswana, using interviewer characteristics such as gender, age and years of experience as IVs.
研究动机与目标
- 解决在健康与社会科学中常见的、因结果缺失非随机(MNAR)而产生的选择偏差问题。
- 通过满足排除限制与相关性条件的工具变量(IV),实现对MNAR下完整数据分布的非参数识别。
- 为受MNAR影响的结果的总体均值,开发一种半参数、双重稳健的估计量,以增强对模型误设的稳健性。
- 通过在博茨瓦纳莫丘迪的HIV血清流行率研究中应用该方法,展示其实际应用价值,其中拒绝检测导致MNAR。
- 在较弱的参数假设下,将现有缺失随机(MAR)方法扩展至处理MNAR问题。
提出的方法
- 提出在IV满足排除限制与相关性条件时,对MNAR下完整数据分布进行非参数识别的必要与充分条件。
- 通过结合逆概率加权与结果回归,开发一种结果均值的双重稳健估计量,确保若任一模型设定正确,估计结果均具一致性。
- 使用估计方程来估计结果、缺失性与IV的联合分布,整合倾向得分与结果回归模型。
- 采用两阶段估计程序:首先估计给定协变量下IV的条件分布,然后估计结果模型与选择偏差函数。
- 将该方法应用于博茨瓦纳一项研究的真实数据,使用访谈员性别、年龄与经验作为IV,以校正HIV检测中不可忽略的非响应偏差。
- 实施数值优化与根求解算法(如uniroot),以求解选择偏差参数并估计MNAR下的均值。
实验结果
研究问题
- RQ1在数据缺失非随机(MNAR)且存在工具变量(IV)的情况下,完整数据分布在何种条件下可实现非参数识别?
- RQ2如何构建一种在MNAR下对总体均值的双重稳健估计量,以确保若倾向得分或结果模型任一正确设定,估计结果均具一致性?
- RQ3所提出的方法能否有效校正现实世界健康调查中因不可忽略的非响应导致的选择偏差,例如HIV血清流行率研究?
- RQ4在实证应用中,哪些实际可验证的识别条件既具有理论有效性又具备可行性?
- RQ5在MNAR条件下,所提出估计量的表现与标准方法(如完整案例分析或基于MAR的方法)相比如何?
主要发现
- 本文建立了在使用工具变量时,对MNAR下完整数据分布进行非参数识别的必要与充分条件,拓展了因果推断领域的既有研究。
- 提出一种新颖的双重稳健估计量用于结果均值,即使倾向得分或结果模型中任一设定正确,估计结果仍具一致性,显著增强了对模型误设的稳健性。
- 该方法在博茨瓦纳莫丘迪的HIV血清流行率研究中成功校正了选择偏差,其中拒绝检测具有不可忽略性,且与访谈员特征相关。
- 模拟与实证结果表明,当数据为MNAR时,所提估计量优于完整案例分析与标准MAR方法。
- 与完全参数模型相比,该估计量在更弱的假设下实现有效推断,降低了对模型误设的敏感性。
- 实证应用表明,访谈员特征(性别、年龄、经验)可作为有效工具变量,使MNAR下的识别与估计成为可能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。