[论文解读] Integrating Testing and Operation-related Quantitative Evidences in Assurance Cases to Argue Safety of Data-Driven AI/ML Components
本文提出了一种整体性的保证案例框架,通过定量整合测试结果、运行时操作数据、范围合规性以及测试数据质量,严格论证数据驱动的AI/ML组件的安全性。通过将这些因素共同数学建模,该框架使安全主张更具说服力,基于证据,优于仅依赖统计测试失败率的主张。
In the future, AI will increasingly find its way into systems that can potentially cause physical harm to humans. For such safety-critical systems, it must be demonstrated that their residual risk does not exceed what is acceptable. This includes, in particular, the AI components that are part of such systems' safety-related functions. Assurance cases are an intensively discussed option today for specifying a sound and comprehensive safety argument to demonstrate a system's safety. In previous work, it has been suggested to argue safety for AI components by structuring assurance cases based on two complementary risk acceptance criteria. One of these criteria is used to derive quantitative targets regarding the AI. The argumentation structures commonly proposed to show the achievement of such quantitative targets, however, focus on failure rates from statistical testing. Further important aspects are only considered in a qualitative manner -- if at all. In contrast, this paper proposes a more holistic argumentation structure for having achieved the target, namely a structure that integrates test results with runtime aspects and the impact of scope compliance and test data quality in a quantitative manner. We elaborate different argumentation options, present the underlying mathematical considerations, and discuss resulting implications for their practical application. Using the proposed argumentation structure might not only increase the integrity of assurance cases but may also allow claims on quantitative targets that would not be justifiable otherwise.
研究动机与目标
- 解决现有保证案例中对运行时与测试证据仅作定性处理而非定量处理的缺陷。
- 开发一种统一的论证结构,将测试结果、运行时行为、数据质量与范围合规性在安全关键场景中整合。
- 为安全关键系统中的AI/ML组件实现更具说服力和可辩护的定量安全目标。
- 通过结合真实世界运行证据与受控测试,提升保证案例的完整性与可信度。
提出的方法
- 提出一种正式的论证结构,整合多种定量证据来源:统计测试结果、运行时故障率、测试数据质量度量指标以及范围合规水平。
- 引入数学模型,将这些证据类型整合为一致的安全论证,使用概率推理评估残余风险。
- 应用同时考虑基于测试的故障率与真实世界条件下运行性能的风险接受标准。
- 采用分层保证案例结构,使每个组件(测试、运行、数据质量)均以定量方式为整体安全主张做出贡献。
- 采用一种支持对证据输入进行敏感性分析的框架,提升安全论证的透明度与可审计性。
实验结果
研究问题
- RQ1如何在保证案例中正式结合测试与运行时证据,以支持AI/ML组件的安全主张?
- RQ2何种数学模型可将测试数据质量与范围合规性整合进定量安全论证?
- RQ3与仅依赖测试的方法相比,包含运行时数据在多大程度上提升了安全主张的可辩护性?
- RQ4如何构建保证案例,以在整合异构证据来源时保持其完整性?
主要发现
- 所提出的框架通过定量整合测试结果、运行时数据、数据质量与范围合规性,使安全主张更具说服力。
- 引入运行时证据可支持仅凭测试数据无法成立的安全主张,尤其在分布外场景下更为显著。
- 数据质量与范围合规性度量的整合,使残余风险估计比孤立测试更为准确。
- 数学建模方法支持透明度与可审计性,增强了对保证案例的信任。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。