Skip to main content
QUICK REVIEW

[论文解读] A Survey on the Evaluation of Clone Detection Performance and Benchmarking

Jeffrey Svajlenko, Chanchal K. Roy|arXiv (Cornell University)|Jun 28, 2020
Software Engineering Research参考文献 257被引用 10
一句话总结

本文全面综述了软件工程领域克隆检测工具的评估、基准测试实践及实证验证方法。研究分析了184种克隆检测工具,基于召回率、精确率、执行时间和可扩展性评估其性能,并指出现有评估实践中的关键缺陷,倡导采用标准化基准和严谨的实证验证,以提升克隆检测研究的质量与可复现性。

ABSTRACT

There are a great many clone detection tools proposed in the literature. In this paper, we investigate the state of clone detection tool evaluation. We begin by surveying the clone detection benchmarks, and performing a multi-faceted evaluation and comparison of their features and capabilities. We then survey the existing clone detection tool and technique publications, and evaluate how the authors of these works evaluate their own tools/techniques. We rank the individual works by how well they measure recall, precision, execution time and scalability. We select the works the best evaluate all four metrics as exemplars that should be considered by future researchers publishing clone detection tools/techniques when designing the empirical evaluation of their tool/technique. We measure statistics on tool evaluation by the authors, and find that evaluation is poor amongst the authors. We finish our investigation into clone detection evaluation by surveying the existing tool comparison studies, including both the qualitative and quantitative studies.

研究动机与目标

  • 调查软件工程研究中克隆检测工具评估与基准测试的现状。
  • 识别作者在测量召回率、精确率、执行时间和可扩展性方面评估克隆检测工具时存在的缺陷。
  • 评估现有克隆检测基准的质量、完整性和作为标准化工具比较基础的适用性。
  • 分析工具比较研究,识别高质量、严谨的评估案例,以供未来研究作为范例。
  • 通过识别最佳实践并倡导在克隆检测研究中采用一致、全面的实证验证,推动评估标准的改进。

提出的方法

  • 系统性地调研了截至2017年发表文献中的184种克隆检测工具与技术。
  • 基于四个关键性能指标(召回率、精确率、执行时间与可扩展性)评估每种工具。
  • 分析基准测试中使用的参考语料库与参考机制,以评估评估结果的有效性与可靠性。
  • 根据设计、范围与评估标准,对10个主要克隆检测基准(如Bellon’s、BigCloneBench、ForkSim)进行分类与比较。
  • 调研工具作者的评估实践,包括指标使用频率、评估范围以及各指标之间的相关性。
  • 识别出评估全面、多指标的典范研究,作为未来研究与同行评审的参考范例。

实验结果

研究问题

  • RQ1当前用于克隆检测的基准有哪些?它们在设计、覆盖范围与质量方面如何比较?
  • RQ2克隆检测工具作者当前如何评估其工具的召回率、精确率、执行时间和可扩展性?
  • RQ3克隆检测研究出版物中的评估实践在多大程度上实现了标准化或严谨性?
  • RQ4克隆检测评估与工具比较研究面临的主要有效性威胁是什么?
  • RQ5哪些研究可作为克隆检测工具高质量、全面评估的典范?

主要发现

  • 仅有不到10%的克隆检测工具论文在召回率、精确率、执行时间和可扩展性四项关键指标上进行了全面评估。
  • 许多工具作者依赖非正式或不完整的评估,常省略精确率或可扩展性测量,从而削弱了其宣称的性能改进的可靠性。
  • 最严谨的评估出现在使用BigCloneBench和Bellon’s Benchmark的研究中,这些基准提供了结构良好、经人工验证的参考语料库。
  • 标准化基准严重缺乏,这导致历史上新工具的接受更多基于算法新颖性,而非实证性能表现。
  • 仅有少数高质量的工具比较研究存在,但它们证明了多指标、可复现评估的可行性与必要性,是可信研究的关键。
  • 典范研究(如Falke等人与Qu等人所做)展示了使用多指标与参考语料库的全面评估,为未来工作设定了基准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。