[论文解读] Pooling multiple imputations when the sample happens to be the population
本文提出了一种在样本为总体时的简化多重插补汇总规则,消除了推断中的抽样方差。通过仅关注缺失数据带来的变异,该方法产生的置信区间更短且覆盖正确,相较于传统Rubin规则(后者在有限总体设定中过度估计方差,降低统计功效)。
Current pooling rules for multiply imputed data assume infinite populations. In some situations this assumption is not feasible as every unit in the population has been observed, potentially leading to over-covered population estimates. We simplify the existing pooling rules for situations where the sampling variance is not of interest. We compare these rules to the conventional pooling rules and demonstrate their use in a situation where there is no sampling variance. Using the standard pooling rules in situations where sampling variance should not be considered, leads to overestimation of the variance of the estimates of interest, especially when the amount of missingness is not very large. As a result, populations estimates are over-covered, which may lead to a loss of statistical power. We conclude that the theory of multiple imputation can be extended to the situation where the sample happens to be the population. The simplified pooling rules can be easily implemented to obtain valid inference in cases where we have observed essentially all units and in simulation studies addressing the missingness mechanism only.
研究动机与目标
- 解决在所有总体单位均被观测时多重插补中对方差的过度估计问题,从而导致置信区间过度覆盖。
- 开发并验证排除抽样方差的简化汇总规则,仅关注缺失数据机制带来的变异。
- 在抽样方差无关的场景(如完整登记册或针对缺失数据机制的模拟研究)中,提高统计效率与功效。
- 为有限总体情境下提供一种实用且数学上简洁的替代传统Rubin规则的方法。
提出的方法
- 提出一种修改后的汇总规则:当样本为总体时,将组内插补方差的平均值 $\bar{U}$ 设为零,因为此时不存在抽样方差。
- 使用Rubin的方差分解,但省略 $\bar{U}$,因此总方差变为 $T = B + B/m$,其中 $B$ 为组间插补方差。
- 应用简化规则来估计感兴趣参数的总方差,自由度为 $\nu = m - 1$。
- 推导出由于非响应导致的方差相对增加量为 $r = \infty$,反映抽样方差的缺失。
- 通过模拟验证该方法,使用包含 $N=1000$ 个单位的有限总体,数据为多元正态分布,缺失率在10%至95%之间变化。
- 在10,000次模拟中,比较传统Rubin规则与简化规则在覆盖概率、置信区间宽度和统计功效方面的表现。
实验结果
研究问题
- RQ1当样本为总体时,传统Rubin汇总规则是否因包含无关的抽样方差而过度估计方差,导致置信区间过度覆盖?
- RQ2是否可以通过排除抽样方差的简化汇总规则,在有限总体设定中提高统计功效与精度?
- RQ3在不同缺失水平下,简化汇总规则与传统汇总规则在覆盖概率与区间宽度方面的表现如何比较?
- RQ4该简化规则是否适用于仅关注缺失数据机制、不包含抽样方差的模拟研究?
- RQ5缺失程度如何影响组间插补方差 ($B$) 与组内插补方差 ($\bar{U}$) 在总方差中的相对贡献?
主要发现
- 当所有总体单位均被观测时,传统Rubin规则导致95%置信区间的过度覆盖,覆盖率超过95%(例如在10%缺失率时为1.000),这是由于包含了无关的抽样方差。
- 简化汇总规则在所有缺失率水平下均实现了正确的95%覆盖,覆盖率稳定在约0.95,表明推断有效。
- 与传统规则相比,简化规则下的置信区间显著更窄——例如在10%缺失率时,$Y_1$ 的区间宽度分别为0.06与0.13,从而提高了统计功效。
- 随着缺失率增加,$\bar{U}$ 对总方差的相对贡献减小,且传统规则下的自由度趋近于 $m-1$,与简化规则趋于一致。
- 当所有数据均缺失时,两种汇总方法等价,因为 $\bar{U} = 0$ 且 $r = \infty$,在极端情况下验证了方法的一致性。
- 该简化方法在专注于缺失数据机制的模拟研究中尤为有益,因为它避免了建模抽样方差的复杂性,同时保持了有效的推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。