[论文解读] Testing normality using the summary statistics with application to meta-analysis
本文提出三种新颖的检验统计量,用于仅基于汇总统计量(样本量、中位数、最小值/最大值、四分位距)评估临床数据中的正态性(对称性),在元分析中应用均值与标准差估计方法之前进行检验。这些检验旨在检测偏态数据,通过在转换前排除非正态研究,提高Cohen's d和Hedges' g等效应量计算的可靠性,从而减少循证医学中的估计偏差。
As the most important tool to provide high-level evidence-based medicine, researchers can statistically summarize and combine data from multiple studies by conducting meta-analysis. In meta-analysis, mean differences are frequently used effect size measurements to deal with continuous data, such as the Cohen's d statistic and Hedges' g statistic values. To calculate the mean difference based effect sizes, the sample mean and standard deviation are two essential summary measures. However, many of the clinical reports tend not to directly record the sample mean and standard deviation. Instead, the sample size, median, minimum and maximum values and/or the first and third quartiles are reported. As a result, researchers have to transform the reported information to the sample mean and standard deviation for further compute the effect size. Since most of the popular transformation methods were developed upon the normality assumption of the underlying data, it is necessary to perform a pre-test before transforming the summary statistics. In this article, we had introduced test statistics for three popular scenarios in meta-analysis. We suggests medical researchers to perform a normality test of the selected studies before using them to conduct further analysis. Moreover, we applied three different case studies to demonstrate the usage of the newly proposed test statistics. The real data case studies indicate that the new test statistics are easy to apply in practice and by following the recommended path to conduct the meta-analysis, researchers can obtain more reliable conclusions.
研究动机与目标
- 解决在将非正态汇总统计量(如中位数、四分位距)转换为均值与标准差时,元分析中可能出现的偏倚风险。
- 开发一种仅使用常见报告的汇总统计量(样本量、中位数、最小值/最大值、四分位距)的正态性预检方法。
- 通过在转换前筛选出具有偏态基础数据的研究,提高元分析中效应量估计的可靠性。
- 为研究人员提供一种实用且统计上可靠的程序,用于在应用Wan等(2014)和Luo等(2017)等现有估计方法前评估数据对称性。
提出的方法
- 提出三种基于不同汇总测量组合的检验统计量:最小值/最大值(T₁)、第一/第三四分位数(T₂)以及两者结合(T₃)。
- 推导出调整样本量和预期正态顺序统计量的系数函数 τ(n)、φ(n) 和 κ(n),使用标准正态分布的分位函数 Φ⁻¹。
- 将对称性偏离的标准化度量(如 (a + b - 2m)/(b - a))作为可检验的偏度度量,并通过依赖样本量的系数进行缩放。
- 应用渐近理论,确保在对称性原假设下,检验统计量服从标准正态分布。
- 将检验整合到推荐的工作流程中:预检对称性 → 排除偏态研究 → 估计均值与标准差 → 计算效应量。
- 通过模拟研究和三个真实世界元分析验证该方法,证明其在效应量估计中具有更高的准确性和一致性。
实验结果
研究问题
- RQ1能否仅基于汇总统计量(样本量、中位数、最小值/最大值、四分位距)在无原始数据的情况下可靠地构建正态性检验?
- RQ2当仅可获得汇总测量时,所提出的检验统计量在检测临床数据中偏态分布方面的有效性如何?
- RQ3通过所提出的检验识别出非正态的研究并予以排除,是否能带来更准确、更可靠的元分析效应量估计?
- RQ4在元分析环境中,针对不同类型的偏态分布,所提出的检验在统计功效和性能方面如何比较?
主要发现
- 在模拟研究中,所提出的检验统计量(T₁、T₂、T₃)表现出接近1的统计功效,表明其具有强大的检测非正态(偏态)数据的能力。
- 在关于他汀类药物治疗与血浆脂质水平的真实案例研究中,对称性检验导致对总胆固醇和LDL-C水平的结论与原始元分析相比发生反转。
- 在BNP与COPD的元分析中,检验识别出若干研究(如第5项研究,中位数为50,Q3为51)可能为右偏态,提示假设对称性的转换方法可能引入误差。
- 三个真实数据应用的森林图显示,排除偏态研究后,合并效应量估计值发生变化,尤其在MMP-3和TIMP-I中表现明显,表明对称性检验会影响临床结论。
- 检验统计量具有稳健性且易于实际应用,具有闭式表达式,无需原始数据,适用于元分析工作流程中的常规使用。
- 推荐流程——在转换前进行对称性检验——可带来更忠实、更可靠的元分析结果,尤其在使用Wan等(2014)和Luo等(2017)等估计器时更为显著。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。