Skip to main content
QUICK REVIEW

[论文解读] Reliability Assessment and Safety Arguments for Machine Learning Components in Assuring Learning-Enabled Autonomous Systems

Xingyu Zhao, Wei Huang|arXiv (Cornell University)|Nov 30, 2021
Software Reliability and Analysis Research参考文献 68被引用 5
一句话总结

本文提出了一种面向学习型自主系统中机器学习分类器的模型无关可靠性评估模型(RAM),通过整合操作剖面与鲁棒性验证,量化可靠性并支持概率性安全论证。该框架支持组件级别的可靠性声明,并通过自主水下航行器的仿真验证了其有效性。

ABSTRACT

The increasing use of Machine Learning (ML) components embedded in autonomous systems -- so-called Learning-Enabled Systems (LES) -- has resulted in the pressing need to assure their functional safety. As for traditional functional safety, the emerging consensus within both, industry and academia, is to use assurance cases for this purpose. Typically assurance cases support claims of reliability in support of safety, and can be viewed as a structured way of organising arguments and evidence generated from safety analysis and reliability modelling activities. While such assurance activities are traditionally guided by consensus-based standards developed from vast engineering experience, LES pose new challenges in safety-critical application due to the characteristics and design of ML models. In this article, we first present an overall assurance framework for LES with an emphasis on quantitative aspects, e.g., breaking down system-level safety targets to component-level requirements and supporting claims stated in reliability metrics. We then introduce a novel model-agnostic Reliability Assessment Model (RAM) for ML classifiers that utilises the operational profile and robustness verification evidence. We discuss the model assumptions and the inherent challenges of assessing ML reliability uncovered by our RAM and propose practical solutions. Probabilistic safety arguments at the lower ML component-level are also developed based on the RAM. Finally, to evaluate and demonstrate our methods, we not only conduct experiments on synthetic/benchmark datasets but also demonstrate the scope of our methods with a comprehensive case study on Autonomous Underwater Vehicles in simulation.

研究动机与目标

  • 解决学习型自主系统(LES)中功能安全保证的挑战,传统安全方法因机器学习模型特性而失效。
  • 开发一种定量保障框架,将系统级安全目标分解为组件级可靠性需求。
  • 构建一种面向机器学习分类器的模型无关可靠性评估模型(RAM),整合操作剖面数据与鲁棒性验证证据。
  • 基于RAM在机器学习组件层面构建概率性安全论证,以支持安全论证案例的构建。
  • 通过合成数据集与一项关于自主水下航行器的全面案例研究,展示该框架的适用性与有效性。

提出的方法

  • 设计一种模型无关的可靠性评估模型(RAM),利用操作剖面数据与鲁棒性验证结果估算分类器的可靠性。
  • 将操作剖面信息——代表真实世界输入分布——整合到可靠性量化中,以反映实际使用条件。
  • 将鲁棒性验证证据(如对抗性测试或形式化验证)作为关键输入,评估在扰动下的可靠性。
  • 通过结合RAM输出与统计可靠性指标及置信区间,在组件层面构建概率性安全论证。
  • 应用该框架将系统级安全目标分解为可度量的组件级可靠性目标。
  • 通过基准数据集与合成数据集上的实验,以及一项基于仿真的自主水下航行器完整案例研究,验证该方法。

实验结果

研究问题

  • RQ1如何以模型无关且考虑操作剖面的方式,对LES中机器学习分类器的可靠性进行定量评估?
  • RQ2鲁棒性验证在增强机器学习组件可靠性声明可信度方面发挥什么作用?
  • RQ3如何推导并使用组件级可靠性度量,以支持保证案例中的系统级安全声明?
  • RQ4所提出的RAM框架在具有复杂操作剖面的真实世界自主系统中可应用到何种程度?
  • RQ5操作剖面与鲁棒性证据的整合在多大程度上提升了机器学习组件安全论证的可辩护性?

主要发现

  • 所提出的RAM无需依赖模型特定假设,即可实现对机器学习分类器可靠性的可靠估计,显著提升了其在多样化架构中的适用性。
  • 整合操作剖面显著提升了可靠性评估的现实性与相关性,真实反映了实际部署条件。
  • 鲁棒性验证证据在可靠性量化中发挥了实质性作用,尤其在识别输入扰动下的失效模式方面表现突出。
  • 该框架成功支持了在组件层面构建概率性安全论证,实现了可追溯且基于证据的保证案例。
  • 对自主水下航行器的案例研究证明了该框架在复杂、关键安全仿真环境中的可扩展性与实际应用价值。
  • 将定量可靠性度量与安全论证结构相结合,使得学习型系统的安全声明更具可辩护性与可审计性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。