Skip to main content
QUICK REVIEW

[Paper Review] Adapting SQuaRE for Quality Assessment of Artificial Intelligence Systems

Hiroshi Kuwajima, Fuyuki Ishikawa|arXiv (Cornell University)|Jul 31, 2019
Adversarial Robustness in Machine Learning8 references4 citations
TL;DR

This paper adapts the ISO/IEC SQuaRE framework for quality assessment of machine learning-based AI systems by integrating principles from the European Commission's Ethics Guidelines for Trustworthy AI. It identifies key updates to SQuaRE’s quality characteristics and sub-characteristics—particularly expanding reliability to include reproducibility, adding privacy as a sub-characteristic, and introducing accountability at the meta-level—thereby enabling holistic evaluation of AI system quality beyond traditional software metrics.

ABSTRACT

More and more software practitioners are tackling towards industrial applications of artificial intelligence (AI) systems, especially those based on machine learning (ML). However, many of existing principles and approaches to traditional systems do not work effectively for the system behavior obtained by training not by logical design. In addition, unique kinds of requirements are emerging such as fairness and explainability. To provide clear guidance to understand and tackle these difficulties, we present an analysis on what quality concepts we should evaluate for AI systems. We base our discussion on ISO/IEC 25000 series, known as SQuaRE, and identify how it should be adapted for the unique nature of ML and $ extit{Ethics guidelines for trustworthy AI}$ from European Commission. We thus provide holistic insights for quality of AI systems by incorporating the ML nature and AI ethics to the traditional software quality concepts.

Motivation & Objective

  • Address the growing need for quality assessment frameworks tailored to machine learning-based AI systems, which differ fundamentally from traditional software in design and behavior.
  • Identify gaps in the existing SQuaRE framework when applied to AI systems due to the black-box, data-driven nature of ML components.
  • Integrate ethical requirements from the European Commission’s Ethics Guidelines for Trustworthy AI into SQuaRE to enhance its relevance for real-world AI applications.
  • Provide actionable, concept-level updates to SQuaRE’s quality model to support industrial practitioners in evaluating AI system quality holistically.
  • Lay the foundation for future work on internal quality aspects such as data curation, model training, and runtime monitoring in AI development.

Proposed method

  • Conduct a comparative analysis between the ISO/IEC 25000 SQuaRE framework and the European Commission’s Ethics Guidelines for Trustworthy AI.
  • Evaluate each SQuaRE quality characteristic and sub-characteristic for validity and completeness when applied to ML-based AI systems.
  • Identify SQuaRE concepts invalidated by ML-specific traits (e.g., lack of explainability, parameter-driven behavior) and propose necessary modifications.
  • Map ethical requirements from the Ethics Guidelines—especially fairness, explainability, accountability, and privacy—onto SQuaRE’s hierarchical structure.
  • Propose new sub-characteristics (e.g., Privacy under Security, Reproducibility under Reliability) and meta-level considerations (e.g., auditability and traceability for accountability).
  • Use a tree-structured analysis to align ethical sub-requirements with SQuaRE’s quality model, ensuring traceability and extensibility.

Experimental results

Research questions

  • RQ1How does the black-box, data-driven nature of machine learning invalidate or limit the applicability of traditional SQuaRE quality characteristics?
  • RQ2Which quality aspects from the European Commission’s Ethics Guidelines for Trustworthy AI are missing or inadequately represented in the current SQuaRE framework?
  • RQ3What specific modifications are needed to extend SQuaRE’s quality model to include ML-specific and ethical quality dimensions such as fairness, explainability, and reproducibility?
  • RQ4How can accountability and auditability be meaningfully integrated into SQuaRE without overhauling its core structure?
  • RQ5To what extent do existing SQuaRE quality measures support the evaluation of AI system quality beyond accuracy and reliability?

Key findings

  • The current SQuaRE framework does not adequately address the unique quality challenges of ML-based AI systems, such as lack of explainability and sensitivity to training data.
  • The concept of 'reproducibility'—critical in ML due to randomness and hyperparameter sensitivity—is not explicitly covered in SQuaRE and should be added as a sub-characteristic under Reliability.
  • Privacy must be elevated from a sub-characteristic of Security to a first-class quality sub-characteristic in the Product quality model to reflect its growing importance.
  • The Ethics Guidelines’ requirement for 'accountability' is not directly supported by SQuaRE and requires a meta-level consideration, such as audit trails and documentation of decisions.
  • The sub-characteristic 'traceability' in SQuaRE is insufficient for AI systems; it must be reinterpreted to include algorithmic decision logging and provenance tracking for compliance and auditing.
  • Most of the Ethics Guidelines’ seven key requirements are not fully covered by SQuaRE, indicating a significant gap in current software quality standards for trustworthy AI.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.