Skip to main content
QUICK REVIEW

[论文解读] Engineering problems in machine learning systems

Hiroshi Kuwajima, Hirotoshi Yasuoka|arXiv (Cornell University)|Apr 1, 2019
Adversarial Robustness in Machine Learning参考文献 56被引用 5
一句话总结

本文识别并分类了安全关键型机器学习系统中的核心工程挑战——尤其在自动驾驶领域——重点关注需求与设计规范的缺失、可解释性以及鲁棒性问题。它提出了一种框架,通过将测试数据视为需求、训练数据视为设计,将演绎式需求与数据驱动训练相联系,揭示了需求规范不良和鲁棒性不足严重削弱了传统质量模型(如SQuARE)的有效性。

ABSTRACT

Fatal accidents are a major issue hindering the wide acceptance of safety-critical systems that employ machine learning and deep learning models, such as automated driving vehicles. In order to use machine learning in a safety-critical system, it is necessary to demonstrate the safety and security of the system through engineering processes. However, thus far, no such widely accepted engineering concepts or frameworks have been established for these systems. The key to using a machine learning model in a deductively engineered system is decomposing the data-driven training of machine learning models into requirement, design, and verification, particularly for machine learning models used in safety-critical systems. Simultaneously, open problems and relevant technical fields are not organized in a manner that enables researchers to select a theme and work on it. In this study, we identify, classify, and explore the open problems in engineering (safety-critical) machine learning systems --- that is, in terms of requirement, design, and verification of machine learning models and systems --- as well as discuss related works and research directions, using automated driving vehicles as an example. Our results show that machine learning models are characterized by a lack of requirements specification, lack of design specification, lack of interpretability, and lack of robustness. We also perform a gap analysis on a conventional system quality standard SQuARE with the characteristics of machine learning models to study quality models for machine learning systems. We find that a lack of requirements specification and lack of robustness have the greatest impact on conventional quality models.

研究动机与目标

  • 为解决安全关键型机器学习系统(尤其是自动驾驶领域)缺乏标准化工程流程的问题。
  • 识别并分类机器学习模型在需求、设计与验证方面的开放性问题。
  • 分析传统系统质量模型(如SQuARE)为何无法充分捕捉机器学习特有的特性,例如需求缺失与鲁棒性不足。
  • 提出一个概念性框架,通过测试数据与训练数据,将演绎式需求与数据驱动训练相连接。
  • 为未来研究提供指导,推动制定标准化的质量模型与工程实践,以适用于机器学习系统。

提出的方法

  • 提出一种理想化的训练流程,将测试数据作为需求规范的代理,将训练数据作为设计规范。
  • 将V-Model概念应用于机器学习系统,通过数据将需求与验证相映射。
  • 对SQuARE质量模型与机器学习系统特性之间的差距进行分析,识别关键不匹配点。
  • 将开放性问题划分为四类:需求缺失、设计缺失、可解释性与鲁棒性。
  • 分析机器学习系统中的数据质量,区分测试数据质量(归纳性需求)与训练数据质量(设计规范)。
  • 提出未来研究方向,包括分层验证策略以及超越SQuARE的标准化质量模型。

实验结果

研究问题

  • RQ1如何在机器学习系统中,有意义地将演绎式需求与数据驱动训练相连接?
  • RQ2在安全关键系统中,为机器学习模型指定需求与设计所面临的主要工程挑战是什么?
  • RQ3为何传统系统质量模型(如SQuARE)无法充分评估机器学习系统?
  • RQ4测试数据与训练数据的质量如何影响机器学习系统的可靠性与可验证性?
  • RQ5标准化评估机器学习系统所需的关键质量特征与度量指标是什么?

主要发现

  • 机器学习模型存在根本性的正式需求规范缺失,这严重损害了可追溯性与可验证性。
  • 机器学习系统中缺乏设计规范,使得模型开发的一致性与可复现性难以保障。
  • 可解释性与鲁棒性是主要挑战,尤其是在处理罕见或分布外样本时。
  • 缺乏鲁棒性——特别是对不确定性及极低概率事件的处理能力——对传统质量模型(如SQuARE)的影响最大。
  • 需求规范的缺失,尤其是事前条件与功能细节的不完整,显著降低了标准质量模型的有效性。
  • 开发数据(测试与训练数据)必须被正式纳入质量模型,因为其质量直接影响系统的可靠性与验证结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。