Skip to main content
QUICK REVIEW

[论文解读] An Interval Estimation Approach to Sample Selection Bias

Matthew Tudball, Qingyuan Zhao|arXiv (Cornell University)|Jun 24, 2019
Advanced Causal Inference Techniques参考文献 2被引用 4
一句话总结

本文提出了一种计算高效的区间估计方法,用于在样本选择偏差下对总体参数进行估计,采用随机规划方法推导出在选择机制假设最少情况下的有效置信区间。通过整合常见的辅助数据(如响应率和协变量均值),该方法提高了区间估计的精度,在模拟研究和英国生物银行关于教育对收入影响的真实世界研究中表现出稳健性能。

ABSTRACT

A widespread and largely unaddressed challenge in statistics is that non-random participation in study samples can bias the estimation of parameters of interest. To address this problem, we propose a computationally efficient interval estimator for a class of population parameters encompassing population means, ordinary least squares and instrumental variables estimands which makes minimal assumptions about the selection mechanism. Using results from stochastic programming, we derive valid confidence intervals and hypothesis tests based on this estimator. In addition, we demonstrate how to tighten the intervals by incorporating additional constraints based on population-level information commonly available to researchers, such as survey response rates and covariate means. We conduct a comprehensive simulation study to evaluate the finite sample performance of our estimator and conclude with a real data study on the causal effect of education on income in the highly-selected UK Biobank cohort. We are able to demonstrate that our method can produce informative bounds under relatively few population-level auxiliary constraints.

研究动机与目标

  • 解决统计估计中广泛存在的样本选择偏差问题,即非随机参与导致参数估计失真。
  • 开发一种计算高效的区间估计方法,适用于包括总体均值、OLS 和 IV 估计量在内的广泛参数类别。
  • 通过随机规划方法推导出有效的置信区间和假设检验,而无需对选择机制施加强假设。
  • 通过整合常见的总体水平辅助信息(如调查响应率和协变量均值)来提高区间估计的精度。
  • 通过模拟和一项关于教育与收入因果推断的真实世界研究,展示该方法在有限样本下的性能和实际应用价值。

提出的方法

  • 使用随机规划方法建立区间估计问题,以建模选择机制中的不确定性。
  • 通过求解一个凸优化问题来构建感兴趣参数的置信区间,该问题可限制由于非随机抽样导致的偏差。
  • 对选择过程施加最少的结构假设,仅依赖于选择指示变量和观测到的结果。
  • 将辅助约束(如已知的总体响应率和协变量均值)整合进优化框架,以收紧区间边界。
  • 基于所得的区间估计量推导出有效的统计推断(置信区间和假设检验)。
  • 将该方法应用于模拟数据和英国生物银行的真实世界数据,以评估其在现实条件下的性能。

实验结果

研究问题

  • RQ1能否在对选择机制假设最少的前提下,为样本选择偏差下的总体参数开发一种计算高效的区间估计方法?
  • RQ2诸如已知的调查响应率和协变量均值等辅助约束如何影响所得置信区间的精度?
  • RQ3在现实的模拟设置下,所提出的区间估计方法在有限样本下的表现如何?
  • RQ4该方法能否在高度选择性的现实世界队列(如英国生物银行)中,为教育与收入的因果推断提供有信息量的区间估计?

主要发现

  • 所提出的区间估计方法在对选择机制假设最少的情况下,仍能产生有效的置信区间,确保统计有效性。
  • 整合响应率和协变量均值等辅助约束可显著收紧区间边界,从而在数据需求极少的情况下提高估计精度。
  • 该方法在模拟中表现出强大的有限样本性能,在多种选择机制下均保持接近名义水平的覆盖概率。
  • 在英国生物银行的应用中,尽管队列存在高度选择偏差,该方法仍能对教育对收入的因果效应提供有信息量的估计区间。
  • 该方法保持了计算高效和可扩展性,使其在存在选择偏差的观察性研究中具有实际应用价值。
  • 当存在未观测的混杂因素或选择偏差时,该方法仍能提供有界估计和有效的统计推断,从而支持稳健的因果推断。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。