[论文解读] Statistical Analysis of Data Repeatability Measures
本文在单变量和多变量(M)ANOVA模型下,评估了在不同分布假设下数据可重复性度量的统计功效——具体包括可区分性(Disc)、秩和检验、F检验和组内相关系数(ICC)估计。研究发现,Disc在非正态分布和测量次数增加的情况下,始终优于秩和检验与基于ICC的方法,而基于指纹识别的方法表现最差,对称性破坏会降低所有方法的性能。
The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the characteristics of the data which are repeatable -- the aspects of the data that are able to be identified under a duplicated analysis. Conflictingly, the utility of traditional repeatability measures, such as the intraclass correlation coefficient, under these settings is limited. In recent work, novel data repeatability measures have been introduced in the context where a set of subjects are measured twice or more, including: fingerprinting, rank sums, and generalizations of the intraclass correlation coefficient. However, the relationships between, and the best practices among these measures remains largely unknown. In this manuscript, we formalize a novel repeatability measure, discriminability. We show that it is deterministically linked with the correlation coefficient under univariate random effect models, and has desired property of optimal accuracy for inferential tasks using multivariate measurements. Additionally, we overview and systematically compare repeatability statistics using both theoretical results and simulations. We show that the rank sum statistic is deterministically linked to a consistent estimator of discriminability. The power of permutation tests derived from these measures are compared numerically under Gaussian and non-Gaussian settings, with and without simulated batch effects. Motivated by both theoretical and empirical results, we provide methodological recommendations for each benchmark setting to serve as a resource for future analyses. We believe these recommendations will play an important role towards improving repeatability in fields such as functional magnetic resonance imaging, genomics, pharmacology, and more.
研究动机与目标
- 评估在不同分布假设下,可区分性(Disc)、秩和检验、F检验与ICC估计在衡量数据可重复性方面的相对统计功效。
- 研究非正态性(如对数正态误差)与批次效应如何影响可重复性度量的性能。
- 评估增加重复测量次数对不同统计检验功效的影响。
- 检验当受试者间对称性被破坏时,这些度量的稳健性。
- 比较基于指纹识别的方法与其他可重复性度量在统计功效上的差异。
提出的方法
- 模拟具有受试者特异性随机效应和独立误差项的单变量与多变量(M)ANOVA模型。
- 通过10,000次蒙特卡洛模拟,估计在不同样本量(n = 5 至 100)下的第一类错误率与统计功效。
- 采用可区分性(Disc)作为基于组间与组内方差比的可重复性度量。
- 使用置换检验与非参数秩和检验,与Disc和F检验进行比较。
- 将ICC定义为组间方差与总方差的比值,在多变量情形下采用基于迹的ICC。
- 引入非正态误差结构(如对数正态分布),以评估在分布偏离情况下的稳健性。
实验结果
研究问题
- RQ1在单变量与多变量设置下,可区分性(Disc)是否具有高于秩和检验或基于距离的Wilcoxon检验的统计功效?
- RQ2在非正态误差分布下,Disc与F检验及ICC估计的统计功效相比如何?
- RQ3随着重复测量次数的增加,Disc相较于秩和检验或F检验的相对性能如何变化?
- RQ4在高斯(M)ANOVA模型中,批次效应如何影响Disc、秩和检验与F检验的统计功效?
- RQ5在何种条件下,基于指纹识别的可重复性估计方法表现劣于其他方法?
主要发现
- 在单变量与多变量设置下,Disc始终优于秩和检验与基于ICC的方法,尤其在非正态分布下表现更优。
- 在非高斯(对数正态)误差模型中,即使ICC的确定性变换不再成立,Disc仍保持高于秩和检验与ICC估计的统计功效。
- 随着重复测量次数的增加,Disc相对于其他方法的统计功效优势进一步扩大;尽管结合多个秩和检验可降低但无法消除该优势。
- 在批次效应存在时,秩和检验在高斯模型中优于Disc,表明其性能具有情境依赖性。
- 在所有模拟设置下(包括单变量、多变量与非稀疏情形),基于指纹识别的方法始终为最不具统计功效的方法。
- 当受试者间对称性被破坏时,所有方法的统计功效均下降,表明在该条件下可重复性评估存在根本性局限。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。