Skip to main content
QUICK REVIEW

[论文解读] Estimation of population size based on capture recapture designs and evaluation of the estimation reliability

Yue You, Mark van der Laan|arXiv (Cornell University)|May 12, 2021
Census and Population Estimation参考文献 24被引用 4
一句话总结

本文提出了一种基于目标最大似然估计(TMLE)的框架,用于在捕获-再捕获设计下进行总体规模估计,且参数假设最少。该方法解决了模型识别问题、由空捕获模式引起的偏差,以及高维数据问题,通过使用欠平滑的lasso平滑技术,表明估计的可靠性关键取决于正确识别假设,尤其是对数线性模型中不存在高阶交互作用的假设。

ABSTRACT

We propose a modern method to estimate population size based on capture-recapture designs of K samples. The observed data is formulated as a sample of n i.i.d. K-dimensional vectors of binary indicators, where the k-th component of each vector indicates the subject being caught by the k-th sample, such that only subjects with nonzero capture vectors are observed. The target quantity is the unconditional probability of the vector being nonzero across both observed and unobserved subjects. We cover models assuming a single constraint (identification assumption) on the K-dimensional distribution such that the target quantity is identified and the statistical model is unrestricted. We present solutions for linear and non-linear constraints commonly assumed to identify capture-recapture models, including no K-way interaction in linear and log-linear models, independence or conditional independence. We demonstrate that the choice of constraint has a dramatic impact on the value of the estimand, showing that it is crucial that the constraint is known to hold by design. For the commonly assumed constraint of no K-way interaction in a log-linear model, the statistical target parameter is only defined when each of the $2^K - 1$ observable capture patterns is present, and therefore suffers from the curse of dimensionality. We propose a targeted MLE based on undersmoothed lasso model to smooth across the cells while targeting the fit towards the single valued target parameter of interest. For each identification assumption, we provide simulated inference and confidence intervals to assess the performance on the estimator under correct and incorrect identifying assumptions. We apply the proposed method, alongside existing estimators, to estimate prevalence of a parasitic infection using multi-source surveillance data from a region in southwestern China, under the four identification assumptions.

研究动机与目标

  • 开发一种通用且灵活的框架,用于在最小建模假设下,从K样本捕获-再捕获数据中估计总体规模。
  • 通过基于机器学习的平滑技术,解决高维设置下空捕获模式(维度灾难)的挑战。
  • 评估识别假设(尤其是无K重交互作用或独立性)对估计可靠性与偏差的影响。
  • 通过目标最大似然估计(TMLE)实现渐近高效推断,并提供诚实的置信区间。
  • 通过模拟和真实数据,比较在识别假设正确与错误时,所提方法与现有估计器的性能。

提出的方法

  • 将总体规模估计问题表述为一个非参数目标参数:即个体至少在一个样本中被捕获的无条件概率。
  • 应用一种通用的基于约束的识别框架,允许线性和非线性约束,如对数线性或线性模型中的无K重交互作用。
  • 使用目标最大似然估计(TMLE)生成渐近高效估计器,并减少偏差,尤其在模型误设时表现更优。
  • 采用欠平滑的lasso估计器对未观测到的捕获模式进行平滑,以降低方差,同时保持一致性。
  • 为每种识别假设推导出高效影响曲线,以实现有效推断和置信区间。
  • 将该方法应用于中国西南部多源监测数据,基于四种不同的识别假设估计寄生虫感染的流行率。

实验结果

研究问题

  • RQ1在捕获-再捕获研究中,识别假设的选择(如独立性或无K重交互作用)如何影响估计的总体规模?
  • RQ2当识别假设发生模型误设时,不同估计器在总体规模估计中的偏差程度如何?
  • RQ3在存在空单元的高维捕获-再捕获设置中,结合欠平滑lasso的目标最大似然估计能否提高估计效率和覆盖率?
  • RQ4与传统的插补法和参数估计器相比,基于TMLE的估计器在偏差、方差和置信区间覆盖方面表现如何?
  • RQ5当将该方法应用于具有稀疏捕获模式的真实世界多源监测数据时,其经验表现如何?

主要发现

  • 识别假设的选择——尤其是对数线性模型中无K重交互作用的假设——对估计的总体规模有显著影响,错误假设会导致严重偏差。
  • 所提出的结合欠平滑lasso的TMLE估计器在置信区间覆盖方面优于参数模型,尤其是在存在空单元时,因其采用更诚实、更宽的区间,能够反映模型不确定性。
  • 当识别假设被违反时,所有估计器(包括复杂的机器学习方法)均产生偏差,凸显了基于设计的假设验证的必要性。
  • 在无K重交互作用假设下,TMLE估计器比参数替代方法更具鲁棒性,因为它仅做出识别所必需的最小假设。
  • 该方法通过基于lasso的平滑技术,成功缓解了维度灾难,纠正了高维稀疏捕获模式数据中的偏差。
  • 在模拟和真实数据分析中,基于TMLE的估计器在无K重交互作用假设下,相较于插补法、MLE及其他估计器,在偏差减少和覆盖精度方面表现更优。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。