Skip to main content
QUICK REVIEW

[論文レビュー] Integrating Testing and Operation-related Quantitative Evidences in Assurance Cases to Argue Safety of Data-Driven AI/ML Components

Michael Kläs, Lisa Jöckel|arXiv (Cornell University)|Feb 10, 2022
Safety Systems Engineering in Autonomy被引用数 5
ひとこと要約

本論文は、テスト結果、運用ランタイムデータ、範囲適合性、テストデータ品質を定量的に統合することで、データ駆動型AI/MLコンponentの安全性を厳密に主張する包括的保証ケースフレームワークを提案する。これらの要因を数学的に統合することで、単に統計的テスト失敗率に依存するのとは異なり、より説得力があり、証拠に基づいた安全性の主張が可能になる。

ABSTRACT

In the future, AI will increasingly find its way into systems that can potentially cause physical harm to humans. For such safety-critical systems, it must be demonstrated that their residual risk does not exceed what is acceptable. This includes, in particular, the AI components that are part of such systems' safety-related functions. Assurance cases are an intensively discussed option today for specifying a sound and comprehensive safety argument to demonstrate a system's safety. In previous work, it has been suggested to argue safety for AI components by structuring assurance cases based on two complementary risk acceptance criteria. One of these criteria is used to derive quantitative targets regarding the AI. The argumentation structures commonly proposed to show the achievement of such quantitative targets, however, focus on failure rates from statistical testing. Further important aspects are only considered in a qualitative manner -- if at all. In contrast, this paper proposes a more holistic argumentation structure for having achieved the target, namely a structure that integrates test results with runtime aspects and the impact of scope compliance and test data quality in a quantitative manner. We elaborate different argumentation options, present the underlying mathematical considerations, and discuss resulting implications for their practical application. Using the proposed argumentation structure might not only increase the integrity of assurance cases but may also allow claims on quantitative targets that would not be justifiable otherwise.

研究の動機と目的

  • 既存の保証ケースが運用およびテストの証拠を定性的に扱うのに対し、定量的に扱うというギャップを埋める。
  • テスト結果、ランタイム動作、データ品質、範囲適合性を、安全性が重要な文脈で統合する一貫した議論構造を構築すること。
  • 臨界システムにおけるAI/MLコンponentのより説得力があり正当化可能な定量的安全性目標を可能にすること。
  • 制御されたテストに加えて現実世界の運用証拠を組み込むことで、保証ケースの整合性と信頼性を高めること。

提案手法

  • 統計的テスト結果、運用時の故障率、テストデータ品質指標、範囲適合性レベルの複数の定量的証拠源を統合する形式的な議論構造を提案する。
  • 確率的推論を用いて残存リスクを評価するため、これらの証拠タイプを統合する数学的モデルを導入する。
  • テストベースの故障率と現実世界の条件下での運用パフォーマンスの両方を考慮したリスク受容基準を適用する。
  • 各コンponent(テスト、運用、データ品質)が全体の安全性主張に定量的に寄与する階層的な保証ケース構造を採用する。
  • 証拠入力に対する感度分析を可能にするフレームワークを採用し、安全性主張の透明性と監査可能性を向上させる。

実験結果

リサーチクエスチョン

  • RQ1AI/MLコンponentの安全性主張を支援する保証ケースにおいて、テストと運用証拠を形式的にどのように統合できるか。
  • RQ2テストデータ品質と範囲適合性を定量的安全性主張に統合するための数学的モデルは何か。
  • RQ3運用データを組み込むことで、テストのみのアプローチと比較して、安全性主張の説得力がどの程度向上するか。
  • RQ4多様な証拠源を統合する際、保証ケースをどのように構造化すれば整合性を維持できるか。

主な発見

  • 提案されたフレームワークにより、テスト結果、運用データ、データ品質、範囲適合性を定量的に統合することで、より説得力のある安全性主張が可能になる。
  • 運用証拠を組み込むことで、テストデータのみでは正当化できない安全性主張が可能になる。特に分布外のシナリオにおいて顕著である。
  • データ品質と範囲適合性指標の統合により、単独でのテストよりも、より正確な残存リスク推定が可能になる。
  • 数学的モデリングアプローチにより、透明性と監査可能性が向上し、保証ケースに対する信頼が強化される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。