[论文解读] Estimation and imputation in Probabilistic Principal Component Analysis with Missing Not At Random data
本文提出了一种在缺失非随机(MNAR)机制下对概率主成分分析(PPCA)中的缺失数据进行估计与填补的方法,无需建模缺失数据分布。该方法在自掩蔽MNAR机制下建立了PPCA参数的可识别性,并仅基于完整案例数据推导出均值、方差和协方差的一致估计量,从而在低秩模型中实现无偏填补。
Missing Not At Random (MNAR) values lead to significant biases in the data, since the probability of missingness depends on the unobserved values.They are ''not ignorable'' in the sense that they often require defining a model for the missing data mechanism, which makes inference or imputation tasks more complex. Furthermore, this implies a strong extit{a priori} on the parametric form of the distribution.However, some works have obtained guarantees on the estimation of parameters in the presence of MNAR data, without specifying the distribution of missing data \citep{mohan2018estimation, tang2003analysis}. This is very useful in practice, but is limited to simple cases such as self-masked MNAR values in data generated according to linear regression models.We continue this line of research, but extend it to a more general MNAR mechanism, in a more general model of the probabilistic principal component analysis (PPCA), extit{i.e.}, a low-rank model with random effects. We prove identifiability of the PPCA parameters. We then propose an estimation of the loading coefficients and a data imputation method. They are based on estimators of means, variances and covariances of missing variables, for which consistency is discussed. These estimators have the great advantage of being calculated using only the observed data, leveraging the underlying low-rank structure of the data. We illustrate the relevance of the method with numerical experiments on synthetic data and also on real data collected from a medical register.
研究动机与目标
- 解决在数据缺失非随机(MNAR)时PPCA的参数估计与数据填补挑战,此类缺失会引入选择偏差并使推断复杂化。
- 在一类广义的自掩蔽MNAR机制下建立PPCA参数的可识别性,扩展了以往仅限于简单或MAR设定的研究。
- 开发一种方法,无需指定缺失数据机制即可估计MNAR变量的关键矩(均值、方差、协方差),仅依赖于可观测数据。
- 提出一种计算高效、非参数化的低秩模型缺失值填补方法,避免对数据与缺失性进行联合建模的高成本。
提出的方法
- 利用完整案例分析估计MNAR变量的均值、方差和协方差,利用PPCA的低秩结构。
- 通过从PPCA模型中导出的部分线性模型进行代数运算,推导出载荷矩阵和均值参数的估计量。
- 应用图模型与缺失图来建模条件独立性,并在MAR与MNAR机制下推导一致估计量。
- 提出两种策略:一种基于部分线性模型,另一种利用图模型工具识别在缺失存在情况下的可识别关系。
- 实现一种交替算法,通过一致矩估计量交替估计载荷矩阵与填补缺失值。
- 在正则性条件下(假设A1–A9)确保估计量的一致性,即使缺失性依赖于未观测值也成立。
实验结果
研究问题
- RQ1当数据为MNAR时,特别是自掩蔽机制下,PPCA参数是否可识别?
- RQ2能否在不建模缺失数据机制的情况下,推导出MNAR变量均值、方差与协方差的一致估计量?
- RQ3如何利用PPCA的低秩结构,实现在非忽略性缺失数据设置下的估计与填补?
- RQ4在存在MNAR缺失的情况下,什么条件能确保完整案例估计量的矩估计一致性?
- RQ5与最先进方法相比,该方法在估计精度与填补性能方面表现如何?
主要发现
- 本文证明了在一类广义的自掩蔽MNAR机制下PPCA参数的可识别性,将先前结果从MAR或简单MNAR情况扩展至更一般情形。
- 仅基于可观测数据推导出MNAR变量均值、方差与协方差的一致估计量,无需建模缺失分布。
- 在合成数据上,该方法实现了精确的填补与参数估计,其RMSE与相关性恢复性能优于最先进方法。
- 在真实医疗数据(Traumabase®)上,该方法在MNAR机制下表现出稳健性能,偏差低于标准填补技术。
- 该算法计算高效且可扩展,相关代码已公开发布于GitHub以确保可复现性。
- 实证结果证实,该方法在MAR与MNAR机制共存的数据集中,仍能保持估计一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。