[论文解读] A Procedure and Guidelines for Analyzing Groups of Software Engineering Replications
本文提出了一套定制化的分析流程与指南,用于聚合软件工程(SE)复制研究,强调结合使用汇总数据(AD)与个体参与者数据(IPD-S),并采用随机效应模型以提升可靠性与透明度。该方法通过利用原始数据与分层分析,增强了在软件工程中常见的异质性、小样本复制研究中的中介效应检测能力与统计功效。
Context: Researchers from different groups and institutions are collaborating on building groups of experiments by means of replication (i.e., conducting groups of replications). Disparate aggregation techniques are being applied to analyze groups of replications. The application of unsuitable techniques to aggregate replication results may undermine the potential of groups of replications to provide in-depth insights from experiment results. Objectives: Provide an analysis procedure with a set of embedded guidelines to aggregate software engineering (SE) replication results. Method: We compare the characteristics of groups of replications for SE and other mature experimental disciplines such as medicine and pharmacology. In view of their differences, the limitations with regard to the joint data analysis of groups of SE replications and the guidelines provided in mature experimental disciplines to analyze groups of replications, we build an analysis procedure with a set of embedded guidelines specifically tailored to the analysis of groups of SE replications. We apply the proposed analysis procedure to a representative group of SE replications to illustrate its use. Results: All the information contained within the raw data should be leveraged during the aggregation of replication results. The analysis procedure that we propose encourages the use of stratified individual participant data and aggregated data in tandem to analyze groups of SE replications. Conclusion: The aggregation techniques used to analyze groups of replications should be justified in research articles. This will increase the reliability and transparency of joint results. The proposed guidelines should ease this endeavor.
研究动机与目标
- 解决软件工程(SE)复制研究组缺乏标准化、可辩护的聚合技术的问题。
- 识别SE复制研究与医学等成熟学科之间的关键差异,这些差异会影响聚合方法的适用性。
- 开发一套分步分析流程,并嵌入针对SE复制研究组典型局限性的指南。
- 通过推广使用原始数据与效应量报告,提升联合结论的可靠性与透明度。
- 通过补充材料中的R代码、数据集与技术教程,支持可重现性。
提出的方法
- 将医学与药理学中的元分析结构适配至SE情境,同时考虑SE特有的特征,如小样本量与参与者异质性。
- 提出一种双轨分析方法,同步使用汇总数据(AD)与个体参与者数据(IPD-S),以最大化洞察力与统计功效。
- 对AD与IPD-S均采用随机效应模型,以考虑复制研究间的变异性,并提升结果的普遍适用性。
- 使用带交互项的线性混合模型(LMMs)与元回归,以检测个体层面与实验层面的调节因子。
- 对分类调节因子应用子组元分析,并将显著性阈值调整为0.1,以降低探索性分析中的第一类错误风险。
- 强调报告效应量与95%置信区间,而非仅依赖p值,以增强结果的可解释性并减少对统计显著性的误读。
实验结果
研究问题
- RQ1如何聚合软件工程复制研究组,以确保联合结论的可靠与透明?
- RQ2软件工程复制研究与医学等成熟实验学科之间的关键差异是什么?这些差异如何影响聚合技术的选择?
- RQ3在典型SE复制研究组(样本量小、异质性高)中,汇总数据(AD)与个体参与者数据(IPD-S)哪种聚合技术更为合适?
- RQ4如何通过统计建模可靠地识别SE复制研究组中的调节效应?
- RQ5随机效应模型与效应量报告在提升SE复制研究聚合结果有效性方面发挥什么作用?
主要发现
- 同步使用汇总数据(AD)与个体参与者数据(IPD-S)能显著提升SE复制研究联合结论的信息量与可靠性。
- 由于SE复制研究中普遍存在高度变异性与异质性,随机效应模型优于固定效应模型。
- IPD-S结合线性混合模型(LMMs)相比AD,能提供更大的统计灵活性,并更有效地检测个体层面的调节因子。
- 子组元分析与元回归分别在识别实验层面与连续调节因子方面均具有效性,前提是采用适当的显著性阈值。
- 报告效应量与95%置信区间而非依赖p值,可使复制聚合结果更加透明且易于解释。
- 补充的R代码、数据集与教程可实现完全可重现性,并支持所提出分析流程的实际应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。