Skip to main content
QUICK REVIEW

[论文解读] Exploring ML testing in practice -- Lessons learned from an interactive rapid review with Axis Communications

Qunying Song, Markus Borg|arXiv (Cornell University)|Mar 30, 2022
Software Engineering Research被引用 4
一句话总结

本研究通过与Axis Communications的行业从业者及研究人员开展互动式快速综述(IRR),探讨了计算机视觉系统中机器学习(ML)测试的挑战,特别是数据测试方面的挑战。合作成果包括优化的分类体系,识别出12项关键挑战,并提出九条适用于工业场景的技术规则,尽管尚未完全匹配行业需求,但提供了实用且可适应的见解。

ABSTRACT

There is a growing interest in industry and academia in machine learning (ML) testing. We believe that industry and academia need to learn together to produce rigorous and relevant knowledge. In this study, we initiate a collaboration between stakeholders from one case company, one research institute, and one university. To establish a common view of the problem domain, we applied an interactive rapid review of the state of the art. Four researchers from Lund University and RISE Research Institutes and four practitioners from Axis Communications reviewed a set of 180 primary studies on ML testing. We developed a taxonomy for the communication around ML testing challenges and results and identified a list of 12 review questions relevant for Axis Communications. The three most important questions (data testing, metrics for assessment, and test generation) were mapped to the literature, and an in-depth analysis of the 35 primary studies matching the most important question (data testing) was made. A final set of the five best matches were analysed and we reflect on the criteria for applicability and relevance for the industry. The taxonomies are helpful for communication but not final. Furthermore, there was no perfect match to the case company's investigated review question (data testing). However, we extracted relevant approaches from the five studies on a conceptual level to support later context-specific improvements. We found the interactive rapid review approach useful for triggering and aligning communication between the different stakeholders.

研究动机与目标

  • 启动学术界与产业界在真实应用场景中机器学习测试挑战方面的协作研究。
  • 统一Axis Communications从业者与学术研究人员之间的术语与研究重点。
  • 识别并映射与工业特定挑战(尤其是计算机视觉中的数据测试)相关的机器学习测试研究。
  • 从学术文献中提取可迁移的技术规则及适用性标准,以指导工业应用。
  • 支持未来聚焦于安全关键型机器学习系统中数据测试的联合研究与硕士论文项目。

提出的方法

  • 开展了一项涉及8名利益相关方(4名研究人员与4名Axis Communications从业者)的互动式快速综述(IRR)。
  • 采用结构化、协作式流程,通过迭代对齐与讨论,审查了180项关于机器学习测试的原始研究。
  • 基于现有的软件测试与人工智能质量框架,开发并扩展了分类体系(SERP分类体系),用于对挑战与解决方案进行分类。
  • 从Axis提出的挑战中优先筛选出12项实际挑战,其中前三项(数据测试、度量标准、测试生成)已与相关文献建立映射。
  • 对35项关于“如何测试数据集”的研究进行了深入分析,提取出九条技术规则与情境因素。
  • 利用概念契合度、技术可行性与领域一致性等标准,评估了所选研究的相关性与适用性。

实验结果

研究问题

  • RQ1从产业视角来看,哪些研究解决方案在机器学习驱动的计算机视觉系统中数据测试方面最为相关且适用?
  • RQ2研究人员与从业者如何协作识别并优先处理机器学习测试中的关键挑战?
  • RQ3如何从学术文献中提取技术规则与标准,以指导工业界应用机器学习测试?
  • RQ4分类体系如何支持人工智能质量领域中学术研究与工业实践之间的沟通与对齐?
  • RQ5影响机器学习测试技术在真实系统中适用性的关键情境因素有哪些?

主要发现

  • 未发现现有学术研究与Axis Communications特定数据测试挑战之间存在完美匹配,表明在工业计算机视觉系统中,针对具体情境的数据测试仍存在研究空白。
  • 从五项高潜力研究中提取出九条数据测试技术规则,提供了可适配工业需求的概念性框架。
  • 互动式快速综述过程成功促进了学术界与产业界利益相关方之间的共同理解,统一了术语,并建立了信任。
  • 通过整合软件测试与人工智能质量维度扩展后的SERP分类体系,在组织与导航复杂的机器学习测试研究领域方面表现出显著实用性。
  • 从业者认为“意外充分性技术”尤为相关,目前已开始在数据收集流程中应用该方法。
  • 本研究强调了数据质量保障在人工智能质量中的核心地位,尤其是在汽车、航空电子与医疗等安全关键领域。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。