[论文解读] Efficiency Gains from Using Auxiliary Variables in Imputation
本文表明,在多重插补中引入辅助变量(即不参与主要分析但能预测缺失数据的变量)可显著提高估计效率。通过模拟和真实数据研究,作者发现效率提升相当于将样本量增加多达75%,尤其在缺失数据广泛且估计值接近统计显著性时效果更明显。
Imputation models sometimes use auxiliary variables that, though not part of the planned analysis, can improve the accuracy of imputed values and the efficiency of point estimates. A recent article, using evidence from simulations, argued that the use of auxiliary variables in imputation did not improve efficiency. We review the simulation results and find that the use of auxiliary variables did improve efficiency; under some conditions the efficiency gain was equivalent to increasing the sample size by a quarter. We give an example from our own research where the efficiency gained from auxiliary variables was equivalent to increasing the sample size by three quarters, and pushed some estimates from statistical insignificance to significance. For auxiliary variables to make a difference, there must be a lot of missing data, some estimates must be near the border of significance, and the auxiliary variables must be excellent predictors of the missing values.
研究动机与目标
- 挑战‘辅助变量在插补中无效率增益’这一说法。
- 探究在何种条件下辅助变量能显著提高多重插补中的估计精度。
- 量化在现实缺失数据情景下使用辅助变量所带来的效率增益幅度。
- 展示辅助变量对实证研究中统计显著性影响的实际作用。
- 提供证据表明,当缺失数据较严重时,辅助变量可使原本不显著的估计变为显著。
提出的方法
- 作者重新分析了先前研究中声称辅助变量无效率收益的模拟数据。
- 通过比较使用和不使用辅助变量时点估计量的方差,评估效率增益。
- 该方法涉及使用与缺失值相关的辅助变量进行多重插补,随后评估回归系数估计值的精度。
- 作者使用其自身研究中的真实世界数据,说明辅助变量对统计显著性产生的实际影响。
- 他们使用公式量化效率增益:有效样本量增益 = (无辅助变量时的方差 / 有辅助变量时的方差) - 1。
- 分析聚焦于缺失数据比例高且p值接近临界值的场景,此时辅助变量可能使结果从不显著变为显著。
实验结果
研究问题
- RQ1在插补模型中引入辅助变量是否能带来可测量的估计效率增益?
- RQ2在何种条件下,辅助变量的效率增益最为显著?
- RQ3辅助变量能否通过提升精度,将原本统计不显著的估计变为显著?
- RQ4辅助变量带来的效率增益与实际增加样本量的效果相比如何?
- RQ5辅助变量对缺失数据的预测能力强弱,在多大程度上决定了效率提升的幅度?
主要发现
- 在插补中使用辅助变量显著提高了估计效率,与早期认为无益的结论相矛盾。
- 在某些模拟情景中,效率增益等效于样本量增加25%。
- 在一个真实数据示例中,辅助变量带来的效率增益等效于样本量增加75%。
- 当缺失数据广泛且估计值接近统计显著性临界点时,增益最大。
- 对缺失值具有强预测能力的辅助变量带来了最大的效率提升。
- 在一种情况下,引入辅助变量使原本不显著的估计(p > 0.05)变为显著(p < 0.05),证明了其实际影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。