Skip to main content
QUICK REVIEW

[论文解读] A Systematic Mapping Study on Testing of Machine Learning Programs

Salman Sherin, Muhammad Uzair Khan|arXiv (Cornell University)|Jul 11, 2019
Software Testing and Debugging Techniques参考文献 47被引用 8
一句话总结

本篇系统映射研究分析了从1,654篇文献中筛选出的37篇相关论文,以映射机器学习程序测试的现状。研究识别出关键趋势,按机器学习类型和方法对测试技术进行分类,并指出了在实证证据、工具可用性以及非功能性测试方面存在的差距,特别是针对强化学习系统。

ABSTRACT

We aim to conduct a systematic mapping in the area of testing ML programs. We identify, analyze and classify the existing literature to provide an overview of the area. We followed well-established guidelines of systematic mapping to develop a systematic protocol to identify and review the existing literature. We formulate three sets of research questions, define inclusion and exclusion criteria and systematically identify themes for the classification of existing techniques. We also report the quality of the published works using established assessment criteria. we finally selected 37 papers out of 1654 based on our selection criteria up to January 2019. We analyze trends such as contribution facet, research facet, test approach, type of ML and the kind of testing with several other attributes. We also discuss the empirical evidence and reporting quality of selected papers. The data from the study is made publicly available for other researchers and practitioners. We present an overview of the area by answering several research questions. The area is growing rapidly, however, there is lack of enough empirical evidence to compare and assess the effectiveness of the techniques. More publicly available tools are required for use of practitioners and researchers. Further attention is needed on non-functional testing and testing of ML programs using reinforcement learning. We believe that this study can help researchers and practitioners to obtain an overview of the area and identify several sub-areas where more research is required

研究动机与目标

  • 提供对当前机器学习程序测试研究现状的全面概述。
  • 识别并基于机器学习类型、测试方法和研究重点对现有测试技术进行分类。
  • 从实证证据和报告标准的角度评估已发表研究的质量。
  • 突出非功能性测试和基于强化学习的机器学习系统等研究不足的领域。
  • 通过公开数据并识别未来研究方向,支持研究人员和实践者。

提出的方法

  • 采用软件工程领域文献综述的既定指南,开展系统映射研究。
  • 定义了纳入与排除标准,将最初筛选出的1,654篇论文缩减至37篇相关研究,时间截止至2019年1月。
  • 制定了三组研究问题,以指导对所选文献的分类与分析。
  • 根据贡献方面、研究方面、测试方法、机器学习类型和测试类型等属性对论文进行分类。
  • 使用既定的评估标准,对所选论文的质量进行评估,涵盖实证报告与可复现性。
  • 公开发布收集到的数据,以支持未来在机器学习测试领域的研究与工具开发。

实验结果

研究问题

  • RQ1机器学习程序测试领域的主导研究主题与趋势是什么?
  • RQ2测试技术如何按机器学习类型和测试方法进行分类?
  • RQ3已发表的机器学习测试研究中,实证证据与报告质量如何?
  • RQ4哪些机器学习测试领域仍研究不足,特别是非功能性测试和强化学习?
  • RQ5在机器学习测试研究与实践中,工具可用性与可复现性方面存在哪些关键缺口?

主要发现

  • 机器学习程序测试领域发展迅速,但缺乏足够的实证证据来比较或评估不同测试技术的有效性。
  • 在最初筛选出的1,654篇论文中,仅有37篇符合纳入标准,表明文献中存在高度的噪声与无关内容。
  • 目前缺乏公开可用的工具,难以支持研究人员和实践者对机器学习系统进行测试。
  • 非功能性测试以及基于强化学习的机器学习程序测试仍研究不足,是显著的研究缺口。
  • 所选论文的报告质量参差不齐,标准化评估指标和可复现实验设置的使用有限。
  • 本研究的数据集已公开,可供未来在机器学习测试领域的研究、工具开发与基准测试使用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。