Skip to main content
QUICK REVIEW

[论文解读] A Copula-based Imputation Model for Missing Data of Mixed Type in Multilevel Data Sets

Jiali Wang, Bronwyn Loong|arXiv (Cornell University)|Feb 27, 2017
Statistical Methods and Bayesian Inference参考文献 32被引用 5
一句话总结

本文提出了一种基于配对的多重数据插补模型,适用于具有混合变量类型(连续型、有序型、名义型)的多层次数据,并通过随机效应来考虑聚类结构。基于潜在变量框架,结合扩展秩似然与吉布斯抽样,该方法在聚类内相关性较高时,相较于传统方法,显著提升了插补准确性和参数恢复效果。

ABSTRACT

We propose a copula based method to handle missing values in multivariate data of mixed types in multilevel data sets. Building upon the extended rank likelihood of \cite{hoff2007extending} and the multinomial probit model, our model is a latent variable model which is able to capture the relationship among variables of different types as well as accounting for the clustering structure. We fit the model by approximating the posterior distribution of the parameters and the missing values through a Gibbs sampling scheme. We use the multiple imputation procedure to incorporate the uncertainty due to missing values in the analysis of the data. Our proposed method is evaluated through simulations to compare it with several conventional methods of handling missing data. We also apply our method to a data set from a cluster randomized controlled trial of a multidisciplinary intervention in acute stroke units. We conclude that our proposed copula based imputation model for mixed type variables achieves reasonably good imputation accuracy and recovery of parameters in some models of interest, and that adding random effects enhances performance when the clustering effect is strong.

研究动机与目标

  • 解决具有混合变量类型(连续型、有序型、名义型)的多层次数据集中的缺失数据问题,同时考虑聚类结构。
  • 将Hoff(2007)的基于配对的扩展秩似然方法拓展至包含聚类水平相关性的随机效应。
  • 开发一种通过潜在变量框架中的多项式probit模型来处理名义型变量的模型。
  • 通过模拟和一项中风护理试验的真实数据,评估该方法在插补准确性和参数恢复方面的表现。
  • 展示在多层次数据中聚类效应较强时,引入随机效应所带来的优势。

提出的方法

  • 采用潜在变量模型,将观测到的混合类型变量通过标准正态分布的分位数函数(逆正态CDF)转换为标准正态得分。
  • 应用扩展秩似然(Hoff,2007)方法,利用高斯配对结构来建模变量之间的依赖关系。
  • 在聚类层面引入随机效应,采用多元正态分布以捕捉组内相关性。
  • 通过潜在空间中的多项式probit模型对名义型变量进行建模,以支持离散且无序的类别。
  • 采用吉布斯抽样算法来近似参数和缺失值的后验分布。
  • 通过多重插补方法,将缺失数据带来的不确定性传播至后续分析中。

实验结果

研究问题

  • RQ1基于配对的插补模型能否有效处理多层次数据中混合类型变量(连续型、有序型、名义型)?
  • RQ2在存在聚类结构时,引入随机效应如何提升插补准确性?
  • RQ3与传统插补方法(如均值插补、联合建模)相比,该模型在参数恢复方面表现如何?
  • RQ4当变量分布偏离正态分布时,该模型是否仍保持良好性能?
  • RQ5在存在缺失数据的真实世界多层次临床试验中,该模型在恢复治疗效应方面表现如何?

主要发现

  • 基于配对的插补模型在感兴趣模型中实现了合理的插补准确性和真实参数的有效恢复。
  • 当组内相关系数(ICC)较高时,引入随机效应显著提升了性能,尤其是在恢复聚类水平效应方面。
  • 该模型优于传统方法,尤其在非正态分布条件下,表现出对偏离正态性的稳健性。
  • 在QASC试验数据中,该模型成功处理了混合变量类型中高达17%的中度至高比例缺失数据,同时保持了多层次结构。
  • 在多重插补下,对身体和心理健康评分的治疗效应估计稳定且精确,各次插补的p值和置信区间保持一致。
  • 拟合优度分析表明,尽管高斯配对在计算上较为便捷,但在某些情况下可能需要使用其他配对结构(如t-配对)以捕捉尾部依赖性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。