Skip to main content
QUICK REVIEW

[论文解读] Multiple imputation for multilevel data with continuous and binary variables

Vincent Audigier, Ian R. White|TNO Repository|Feb 3, 2017
Statistical Methods and Bayesian Inference参考文献 49被引用 11
一句话总结

本文评估了在存在连续变量和二元变量的多层次数据中,采用多种多重插补(MI)方法,比较了联合模型(JM)与完全条件指定(FCS)方法在系统性缺失与间歇性缺失情况下的表现。研究发现,异方差MI方法优于同方差方法,且方法选择取决于聚类大小与缺失数据模式,其中JM-jomo与FCS-2stage在不同条件下表现优异。

ABSTRACT

We present and compare multiple imputation methods for multilevel continuous and binary data where variables are systematically and sporadically missing. The methods are compared from a theoretical point of view and through an extensive simulation study motivated by a real dataset comprising multiple studies. Simulations are reproducible. The comparisons show why these multiple imputation methods are the most appropriate to handle missing values in a multilevel setting and why their relative performances can vary according to the missing data pattern, the multilevel structure and the type of missing variables. This study shows that valid inferences can only be obtained if the dataset gathers a large number of clusters. In addition, it highlights that heteroscedastic MI methods provide more accurate inferences than homoscedastic methods, which should be reserved for data with few individuals per cluster. Finally, the method of Quartagno and Carpenter (2016a) appears generally accurate for binary variables, the method of Resche-Rigon and White (2016) with large clusters, and the approach of Jolani et al. (2015) with small clusters.

研究动机与目标

  • 评估并比较在系统性缺失与间歇性缺失条件下,适用于包含连续变量与二元变量的多层次数据的多重插补方法。
  • 评估多层次结构、聚类大小与缺失数据模式对插补准确度与推断有效性的影响。
  • 确定在缺失数据为系统性或间歇性缺失的多层次设置下,哪些MI方法能提供有效的统计推断。
  • 在不同条件(包括异方差性与先验分布假设)下,研究联合模型(JM)与完全条件指定(FCS)方法的表现。
  • 基于数据结构、聚类大小与缺失值比例,提供选择MI方法的实际建议。

提出的方法

  • 使用多重插补(MI),在考虑聚类结构的多层次插补模型下生成M个完整数据集。
  • 比较联合模型(JM)方法——特别是JM-jomo,其采用多元正态分布与共轭先验分布——与基于FCS的方法,包括FCS-GLM、FCS-2stage与FCS-2lnorm。
  • 使用具有随机截距的线性混合模型插补连续变量,二元变量则在多层次框架内使用probit或logit链接函数。
  • 应用Rubin规则对多个插补数据集中的估计值进行合并,以确保推断的方差估计正确。
  • 基于GREAT网络的真实多层次数据集开展模拟研究,评估在不同聚类大小、缺失数据模式与变量类型下的性能表现。
  • 采用允许方差分量在聚类间变化的异方差MI模型,相较于同方差模型,可提高准确性。

实验结果

研究问题

  • RQ1在系统性缺失与间歇性缺失条件下,不同多重插补方法在包含连续变量与二元变量的多层次数据中的表现如何?
  • RQ2聚类大小对多层次设置下多重插补的准确度与推断有效性有何影响?
  • RQ3在多层次数据中,异方差与同方差插补模型在覆盖率与偏差方面如何比较?
  • RQ4在不同缺失数据模式下,哪种插补方法——JM-jomo、FCS-2stage、FCS-GLM或FCS-2lnorm——对二元变量的推断最准确?
  • RQ5FCS-based方法在何种情况下会失效?在何种情况下应优先选择联合模型?

主要发现

  • 在多层次数据中,多重插补的可靠推断需要大量聚类;聚类数量过少会导致方差估计偏差与覆盖率下降。
  • 异方差MI方法在聚类水平方差存在差异时,相比同方差方法能提供更准确的推断。
  • JM-jomo方法在聚类数量与不完整二元变量数量较多时,对二元变量的插补总体上具有较高准确性。
  • FCS-2stage在大聚类中表现良好,但当聚类较小时或间歇性缺失值比例较高时应避免使用。
  • FCS-GLM在小聚类中表现优异,尤其当使用probit链接函数而非logit链接函数时。
  • FCS-2lnorm会低估随机系数的变异性,因此在系统性缺失数据集中适用性较差。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。