[论文解读] A Mixed-effects Model for Incomplete Data With Batch-Level Abundance-Dependent Missing-Data Mechanism
本文提出 mixEMM,一种混合效应模型,可联合考虑 iTRAQ 蛋白质组学数据中的批次效应和一种新型的基于批次水平丰度的缺失数据机制(BADMM)。通过使用指数链接函数显式建模缺失概率作为批次水平蛋白丰度的函数,mixEMM 在参数估计准确性和统计功效方面优于传统方法,后者忽略缺失机制或依赖相对丰度。
In mass spectrometry based quantitative proteomics research, the emerging iTRAQ technique has been widely adopted for high throughput protein profiling, as it enables one to measure multiple samples simultaneously in one multiplex experiment and thus greatly enhances the throughput of protein quantification. However, the technical variation across different iTRAQ multiplex experiments is often large due to the dynamic nature of MS instruments. This leads to strong batch effects in the iTRAQ data. Moreover, the iTRAQ data often contain substantial batch-level non-ignorable missingness. Specifically, the abundance measures of a given protein/peptide are often missing altogether in all the samples from the same batch, with the missing probability depending on the combined batch-level abundances. We term this unique missing-data mechanism as the Batch-level Abundance-Dependent Missing-data mechanism (BADMM). We introduce a new method, mixEMM, for analyzing iTRAQ data with batch effects and batch-level non-ignorable missingness. The mixEMM method employs a linear mixed-effects model and explicitly models the batch effects and the BADMM in the likelihood function. With simulation studies, we showed that compared with existing approaches that utilize relative abundances and ignore the missing batches under the missing completely at random assumption, the mixEMM method achieves more accurate parameter estimation and inference.We applied the method to an iTRAQ proteomics data from a breast cancer study and identified phosphopeptides differentially expressed between different breast cancer subtypes. The method can be applied to general clustered data with cluster level non ignorable missing-data mechanisms.
研究动机与目标
- 解决基于 iTRAQ 的定量蛋白质组学实验中批次效应和不可忽略缺失性的问题。
- 建模一种独特的缺失数据机制——基于批次的丰度依赖性缺失数据机制(BADMM),即基于组合的批次水平丰度,整个批次的蛋白测量值可能缺失。
- 通过将 BADMM 显式整合到基于似然的混合效应模型中,改进统计推断,提升估计准确性和统计功效。
- 为具有聚类水平不可忽略缺失性的数据提供可推广的框架,其应用范围可超越蛋白质组学领域。
提出的方法
- 该方法采用线性混合效应模型,以考虑随机批次效应和表型变量的固定效应。
- 使用批次水平平均丰度的指数链接函数来建模批次中缺失的概率,从而捕捉 BADMM。
- 采用期望-条件最大化(ECM)算法计算模型参数的最大似然估计(MLE),以处理缺失数据和随机效应。
- 似然函数对缺失数据机制进行积分,从而在不可忽略缺失性条件下实现有效推断。
- 探索了替代链接函数(如 logit),以增强灵活性,尽管指数函数在计算效率方面更优。
- 该框架可扩展至多变量分析,并可应用于肽段或蛋白水平,支持将肽段水平的结果汇总为蛋白水平的推断。
实验结果
研究问题
- RQ1与忽略缺失性或假设缺失完全随机相比,显式建模 BADMM 在 iTRAQ 蛋白质组学数据中是否能显著提升参数估计的准确性?
- RQ2当批次水平缺失性普遍存在且依赖于丰度时,mixEMM 方法是否能保持统计功效并减少差异表达分析中的偏差?
- RQ3与使用相对丰度并假设缺失可忽略的传统方法相比,mixEMM 的性能如何?
- RQ4使用不同链接函数(如指数函数与 logit 函数)对缺失数据机制进行建模,对估计准确性与计算效率有何影响?
主要发现
- 模拟研究显示,与忽略 BADMM 或依赖相对丰度的传统方法相比,mixEMM 在参数估计准确性和统计功效方面均显著更优。
- 通过正确考虑批次水平缺失性和方差结构,mixEMM 降低了差异表达检验中的第一类和第二类错误率。
- 与 logit 函数相比,BADMM 的指数链接函数在计算成本显著更低的前提下,提供了相当或更优的性能。
- 在 CPTAC 乳腺癌磷酸化蛋白质组学数据集中,mixEMM 成功识别出不同乳腺癌亚型之间的差异表达磷酸肽,灵敏度和精确度均得到提升。
- 该方法在不同缺失率和丰度分布下表现出稳健性,尤其在缺失性与批次水平丰度强相关时表现更优。
- 一个名为 mixEMM 的 R 包将发布于 CRAN,使该方法可广泛应用于具有聚类和不完整结构的类似高通量组学数据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。