Skip to main content
QUICK REVIEW

[论文解读] Comparison of Canonical Correlation and Partial Least Squares analyses of simulated and empirical data

Anthony R. McIntosh|arXiv (Cornell University)|Jul 14, 2021
Sensory Analysis and Statistical Methods被引用 5
一句话总结

本研究在大样本量的模拟数据和真实数据集上,通过抽样方法评估敏感性、可靠性和可重复性,比较了典型相关分析(CCA)与偏最小二乘法(PLS)的表现。结果表明,在弱效应到强效应条件下,CCA与PLS表现相当;然而,当块内相关性较高时,CCA的可靠性下降,而通过在成分得分上预先进行主成分分析(PCA)可缓解此问题;相比之下,PLS即使在小样本量下仍保持良好的可重复性,凸显了在多元分析中同时评估统计显著性与可重复性的必要性。

ABSTRACT

In this paper, we compared the general forms of CCA and PLS on three simulated and two empirical datasets, all having large sample sizes. We took successively smaller subsamples of these data to evaluate sensitivity, reliability, and reproducibility. In null data having no correlation within or between blocks, both methods showed equivalent false positive rates across sample sizes. Both methods also showed equivalent detection in data with weak but reliable effects until sample sizes drop below n=50. In the case of strong effects, both methods showed similar performance unless the correlations of items within one data block were high. For PLS, the results were reproducible across sample sizes for strong effects, except at the smallest sample sizes. On the contrary, the reproducibility for CCA declined when the within-block correlations were high. This was ameliorated if a principal components analysis (PCA) was performed and the component scores used to calculate the cross-block matrix. The outcome of our examination gives three messages. First, for data with reasonable within and between block structure, CCA and PLS give comparable results. Second, if there are high correlations within either block, this can compromise the reliability of CCA results. This known issue of CCA can be remedied with PCA before cross-block calculation. This, however, assumes that the PCA structure is stable for a given sample. Third, null hypothesis testing does not guarantee that the results are reproducible, even with large sample sizes. This final outcome suggests that both statistical significance and reproducibility be assessed for any data.

研究动机与目标

  • 评估典型相关分析(CCA)与偏最小二乘法(PLS)在不同数据条件下的相对表现。
  • 评估CCA与PLS在样本量递减情况下的敏感性、可靠性和可重复性。
  • 研究高块内相关性对CCA结果的影响,并探索相应的解决方法。
  • 检验统计显著性是否足以确保多元分析结果的可重复性。
  • 基于数据结构与样本量,为CCA与PLS的选择提供实用指导。

提出的方法

  • 本研究将CCA与PLS应用于三个模拟数据集和两个真实数据集,样本量均较大。
  • 从每个数据集中逐次抽取子样本,以评估不同样本量下的性能表现。
  • 使用无块内或块间相关性的零假设数据,以评估假阳性率。
  • 对于CCA,先对数据块进行主成分分析(PCA),再计算块间协方差矩阵。
  • 通过比较不同大小子样本的结果,评估可重复性。
  • 独立评估统计显著性与可重复性,以探究二者之间的关系。

实验结果

研究问题

  • RQ1在不同样本量下,CCA与PLS在统计功效与假阳性率方面如何比较?
  • RQ2高块内相关性对CCA结果可靠性的影响,相较于PLS如何?
  • RQ3能否通过PCA预处理提升高块内相关性条件下CCA的可靠性?
  • RQ4CCA或PLS中的统计显著性是否能保证结果在子样本间的可重复性?
  • RQ5在何种条件下,CCA与PLS在真实数据与模拟数据中表现相近?

主要发现

  • 在无相关性的零假设数据中,CCA与PLS在所有样本量下均表现出相当的假阳性率。
  • 对于弱但可靠的效应,两种方法在样本量降至n=50以下前表现相似。
  • 在强效应条件下,除非块内相关性较高,否则CCA与PLS表现相近;当块内相关性高时,CCA的可靠性下降。
  • PLS在各种样本量下均保持高可重复性,即使在小样本量下亦然,仅在最小样本量时略有下降。
  • CCA在高块内相关性下可重复性下降,但若在计算块间矩阵前对成分得分进行PCA预处理,该问题可得到缓解。
  • 研究发现,即使样本量较大,统计显著性也无法确保可重复性,凸显了同时评估这两项指标的必要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。