Skip to main content
QUICK REVIEW

[论文解读] Data Integration by combining big data and survey sample data for finite population inference

Jaekwang Kim, Siu‐Ming Tam|arXiv (Cornell University)|Mar 26, 2020
Survey Methodology and Nonresponse参考文献 24被引用 7
一句话总结

该论文提出了一种数据整合方法,结合大数据与调查样本数据,实现在不依赖缺失随机(MAR)假设条件下的有效有限总体推断。通过将大数据视为存在测量误差的有限总体,并利用校准加权与非参数分类方法校正选择偏差与误分类问题,该方法在一般选择机制下实现了具有一致性的估计。

ABSTRACT

The statistical challenges in using big data for making valid statistical inference in the finite population have been well documented in literature. These challenges are due primarily to statistical bias arising from under-coverage in the big data source to represent the population of interest and measurement errors in the variables available in the data set. By stratifying the population into a big data stratum and a missing data stratum, we can estimate the missing data stratum by using a fully responding probability sample, and hence the population as a whole by using a data integration estimator. By expressing the data integration estimator as a regression estimator, we can handle measurement errors in the variables in big data and also in the probability sample. We also propose a fully nonparametric classification method for identifying the overlapping units and develop a bias-corrected data integration estimator under misclassification errors. Finally, we develop a two-step regression data integration estimator to deal with measurement errors in the probability sample. An advantage of the approach advocated in this paper is that we do not have to make unrealistic missing-at-random assumptions for the methods to work. The proposed method is applied to the real data example using 2015-16 Australian Agricultural Census data.

研究动机与目标

  • 解决在估计有限总体参数时,大数据源中存在的选择偏差与测量误差问题。
  • 克服现有加权与插补方法中普遍采用的缺失随机(MAR)假设的局限性。
  • 开发一种数据整合框架,利用完全响应概率样本估计缺失数据层,从而改善对整个有限总体的推断。
  • 通过回归校准与两步估计方法,处理大数据与调查样本变量中的测量误差。
  • 提出一种非参数分类方法,用于在无法进行精确匹配时识别大数据与调查样本之间的重叠单位。

提出的方法

  • 将有限总体划分为大数据层(存在覆盖问题与测量误差)和缺失数据层(通过概率样本估计)。
  • 将数据整合估计量表示为回归估计量,以校正两个数据源中的测量误差。
  • 采用半监督非参数分类方法,估计倾向得分,以识别大数据与调查样本之间的重叠单位。
  • 提出一种偏差校正的数据整合估计量,以考虑重叠单位识别过程中误分类误差的影响。
  • 开发一种两步回归估计量,以处理概率样本中的测量误差,提升估计的稳健性与效率。
  • 利用大数据源的辅助信息实施校准加权,以提高估计效率并减少偏差。

实验结果

研究问题

  • RQ1当大数据源存在覆盖不足与测量误差时,如何实现有效的有限总体推断?
  • RQ2能否开发一种不依赖缺失随机(MAR)假设的数据整合方法,以校正选择偏差?
  • RQ3当无法进行精确匹配时,如何有效识别大数据与调查样本之间的重叠单位?
  • RQ4如何在统一的数据整合框架内校正大数据与调查样本中的测量误差?
  • RQ5重叠单位识别中的误分类对最终估计量的偏差与方差有何影响?

主要发现

  • 所提方法在无需MAR假设条件下,于一般选择机制下实现了具有一致性的估计,而该假设在实践中往往不可验证。
  • 非参数分类方法成功识别了大数据与调查样本之间的重叠单位,即使在无法进行精确匹配的情况下亦有效。
  • 偏差校正的数据整合估计量有效降低了因重叠单位识别中误分类导致的偏差。
  • 当概率样本中存在测量误差时,两步回归估计量显著提升了估计的效率与稳健性。
  • 通过泰勒线性化与基于校准的设计推断方法实现方差估计,为质量插补与基于校准的估计量均推导出一致的方差估计量。
  • 在2015-16年澳大利亚农业普查数据上的应用表明,该方法在官方统计生产中具有实际应用价值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。