[Paper Review] A Mixed-effects Model for Incomplete Data With Batch-Level Abundance-Dependent Missing-Data Mechanism
This paper proposes mixEMM, a mixed-effects model that jointly accounts for batch effects and a novel batch-level abundance-dependent missing-data mechanism (BADMM) in iTRAQ proteomics data. By explicitly modeling missingness probability as a function of batch-level protein abundances using an exponential link function, mixEMM improves parameter estimation accuracy and statistical power over conventional methods that ignore missingness mechanisms or rely on relative abundances.
In mass spectrometry based quantitative proteomics research, the emerging iTRAQ technique has been widely adopted for high throughput protein profiling, as it enables one to measure multiple samples simultaneously in one multiplex experiment and thus greatly enhances the throughput of protein quantification. However, the technical variation across different iTRAQ multiplex experiments is often large due to the dynamic nature of MS instruments. This leads to strong batch effects in the iTRAQ data. Moreover, the iTRAQ data often contain substantial batch-level non-ignorable missingness. Specifically, the abundance measures of a given protein/peptide are often missing altogether in all the samples from the same batch, with the missing probability depending on the combined batch-level abundances. We term this unique missing-data mechanism as the Batch-level Abundance-Dependent Missing-data mechanism (BADMM). We introduce a new method, mixEMM, for analyzing iTRAQ data with batch effects and batch-level non-ignorable missingness. The mixEMM method employs a linear mixed-effects model and explicitly models the batch effects and the BADMM in the likelihood function. With simulation studies, we showed that compared with existing approaches that utilize relative abundances and ignore the missing batches under the missing completely at random assumption, the mixEMM method achieves more accurate parameter estimation and inference.We applied the method to an iTRAQ proteomics data from a breast cancer study and identified phosphopeptides differentially expressed between different breast cancer subtypes. The method can be applied to general clustered data with cluster level non ignorable missing-data mechanisms.
Motivation & Objective
- To address the challenge of batch effects and non-ignorable missingness in iTRAQ-based quantitative proteomics experiments.
- To model a unique missing-data mechanism—Batch-level Abundance-Dependent Missing-data Mechanism (BADMM)—where entire batches of protein measurements are missing based on combined batch-level abundances.
- To improve statistical inference by explicitly incorporating the BADMM into a likelihood-based mixed-effects model, enhancing estimation accuracy and power.
- To provide a generalizable framework for clustered data with cluster-level non-ignorable missingness, applicable beyond proteomics.
Proposed method
- The method employs a linear mixed-effects model to account for random batch effects and fixed effects of phenotypic variables.
- It models the probability of missingness in a batch using an exponential link function of the batch-level mean abundance, capturing the BADMM.
- The Expectation-Conditional Maximization (ECM) algorithm is used to compute maximum likelihood estimates (MLEs) of model parameters, handling both missing data and random effects.
- The likelihood function integrates over the missing data mechanism, allowing for valid inference under non-ignorable missingness.
- Alternative link functions, such as logit, are explored for flexibility, though the exponential function is preferred for computational efficiency.
- The framework is extendable to multivariate analysis and can be applied at the peptide or protein level, with options for summarizing peptide-level results into protein-level inferences.
Experimental results
Research questions
- RQ1How does explicitly modeling the BADMM improve parameter estimation accuracy in iTRAQ proteomics data compared to ignoring missingness or assuming missing-completely-at-random?
- RQ2Can the mixEMM method maintain statistical power and reduce bias in differential expression analysis when batch-level missingness is abundant and dependent on abundance?
- RQ3How does the performance of mixEMM compare to conventional approaches that use relative abundances and assume ignorable missingness?
- RQ4What is the impact of using different link functions (e.g., exponential vs. logit) for modeling the missing-data mechanism on estimation accuracy and computational efficiency?
Key findings
- Simulation studies show that mixEMM achieves significantly more accurate parameter estimation and better statistical power than conventional methods that ignore the BADMM or rely on relative abundances.
- The mixEMM method reduces type I and type II error rates in differential expression testing by properly accounting for batch-level missingness and variance structure.
- The exponential link function for BADMM provides comparable or better performance than the logit function with substantially lower computational cost.
- In the CPTAC breast cancer phosphoproteomics dataset, mixEMM identified differentially expressed phosphopeptides between breast cancer subtypes with improved sensitivity and precision.
- The method demonstrated robustness across varying missing rates and abundance distributions, particularly when missingness was strongly dependent on batch-level abundance.
- An R package named mixEMM will be available on CRAN, enabling broad application of the method to similar high-throughput omics data with clustered, incomplete structures.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.