[Paper Review] Towards a power analysis for PLS-based methods
This paper introduces a novel Monte Carlo simulation-based power analysis framework for Partial Least Squares (PLS)-based methods, explicitly preserving the latent structure and correlation structure from pilot data to estimate statistical power and optimal sample size. The method uses PLS-derived weights and scores to simulate multivariate data under the alternative hypothesis, enabling accurate power estimation for classification tasks in omics studies.
In recent years, power analysis has become widely used in applied sciences, with the increasing importance of the replicability issue. When distribution-free methods, such as Partial Least Squares (PLS)-based approaches, are considered, formulating power analysis turns out to be challenging. In this study, we introduce the methodological framework of a new procedure for performing power analysis when PLS-based methods are used. Data are simulated by the Monte Carlo method, assuming the null hypothesis of no effect is false and exploiting the latent structure estimated by PLS in the pilot data. In this way, the complex correlation data structure is explicitly considered in power analysis and sample size estimation. The paper offers insights into selecting statistical tests for the power analysis procedure, comparing accuracy-based tests and those based on continuous parameters estimated by PLS. Simulated and real datasets are investigated to show how the method works in practice.
Motivation & Objective
- Address the lack of reliable power analysis methods for distribution-free, likelihood-free PLS-based approaches in multivariate data analysis.
- Overcome the limitations of traditional power analysis in high-dimensional, correlated, and low-sample-size settings common in -omics research.
- Develop a simulation-based framework that respects the true data correlation structure and latent variable structure from PLS models.
- Compare different statistical tests—accuracy-based and continuous parameter-based—for use in power analysis under PLS frameworks.
- Provide a transparent, accessible R package implementation to support reproducibility and adoption in applied research.
Proposed method
- Simulate artificial datasets using Monte Carlo methods, preserving the latent structure (weights and scores) estimated by PLS from pilot data.
- Generate new datasets with the same covariance structure as the pilot data but varying sample sizes to assess power across designs.
- Apply PLS to the simulated datasets to estimate model parameters and evaluate test statistics under the alternative hypothesis.
- Use both accuracy-based tests (e.g., classification error rate) and continuous parameter-based tests (e.g., latent variable loadings) in the power analysis procedure.
- Employ permutation tests to assess significance and control Type I error rates in the simulation framework.
- Implement the method as an R package to ensure transparency, reproducibility, and accessibility for researchers.
Experimental results
Research questions
- RQ1How can power analysis be reliably performed for PLS-based methods when traditional parametric assumptions do not hold?
- RQ2What is the impact of preserving the latent structure and correlation structure from pilot data on the accuracy of power estimation in PLS models?
- RQ3How do accuracy-based tests compare to continuous parameter-based tests in terms of statistical power and Type I error control in PLS-based classification?
- RQ4What sample size is required to achieve adequate statistical power in a case-control PLS-DA study under realistic data correlation structures?
- RQ5Can a simulation-based approach using PLS-derived components outperform existing methods in sample size estimation for high-dimensional omics data?
Key findings
- The proposed method successfully preserves the complex correlation structure and latent variable relationships from pilot data in simulated datasets, enabling realistic power estimation.
- Power estimation using the PLS-based simulation framework shows that sample size requirements are highly sensitive to the number of latent components (A) and effect size in the data.
- The method achieves stable power estimates with 100 Monte Carlo simulations and 200 permutations, with 95% confidence intervals providing reliable precision.
- Accuracy-based tests demonstrated lower statistical power compared to continuous parameter-based tests, especially in low-sample-size scenarios.
- The simulation framework correctly identified the optimal sample size for achieving 80% power in a case-control PLS-DA setting with A=3 components.
- The R package implementation enables researchers to replicate and apply the method to real-world omics datasets with confidence and transparency.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.