[Paper Review] Quality Assurance for Artificial Intelligence: A Study of Industrial Concerns, Challenges and Best Practices
This study investigates industrial concerns, challenges, and best practices in Quality Assurance for Artificial Intelligence (QA4AI) through interviews with 15 practitioners and a survey of 50, identifying correctness as the top priority, followed by model relevance, efficiency, and deployability. It proposes 21 QA4AI practices, with 10 well-supported and 8 marginally agreed upon, offering a practical checklist for industry implementation.
Quality Assurance (QA) aims to prevent mistakes and defects in manufactured products and avoid problems when delivering products or services to customers. QA for AI systems, however, poses particular challenges, given their data-driven and non-deterministic nature as well as more complex architectures and algorithms. While there is growing empirical evidence about practices of machine learning in industrial contexts, little is known about the challenges and best practices of quality assurance for AI systems (QA4AI). In this paper, we report on a mixed-method study of QA4AI in industry practice from various countries and companies. Through interviews with fifteen industry practitioners and a validation survey with 50 practitioner responses, we studied the concerns as well as challenges and best practices in ensuring the QA4AI properties reported in the literature, such as correctness, fairness, interpretability and others. Our findings suggest correctness as the most important property, followed by model relevance, efficiency and deployability. In contrast, transferability (applying knowledge learned in one task to another task), security and fairness are not paid much attention by practitioners compared to other properties. Challenges and solutions are identified for each QA4AI property. For example, interviewees highlighted the trade-off challenge among latency, cost and accuracy for efficiency (latency and cost are parts of efficiency concern). Solutions like model compression are proposed. We identified 21 QA4AI practices across each stage of AI development, with 10 practices being well recognized and another 8 practices being marginally agreed by the survey practitioners.
Motivation & Objective
- To understand industrial perceptions of QA4AI properties such as correctness, fairness, and interpretability.
- To identify key challenges and solutions for each QA4AI property in real-world AI development.
- To extract and validate best practices for QA4AI across the AI development lifecycle.
- To bridge the gap between academic research and industrial application in AI quality assurance.
Proposed method
- Conducted semi-structured interviews with 15 AI practitioners from diverse companies and countries.
- Administered a validation survey to 50 additional industry practitioners to assess consensus on findings.
- Mapped QA4AI properties to a 9-stage AI development workflow to ensure comprehensive coverage.
- Identified 21 QA4AI practices through thematic analysis of interview data and survey responses.
- Classified practices as 'well-supported' (10) or 'marginally agreed' (8) based on survey consensus.
- Analyzed trade-offs (e.g., latency vs. cost vs. accuracy) and tool usage reported by practitioners.
Experimental results
Research questions
- RQ1How do industry practitioners rank the importance of different QA4AI properties?
- RQ2What are the key challenges and solutions for ensuring each QA4AI property in practice?
- RQ3What best practices are recognized and adopted across the AI development lifecycle?
- RQ4How do practitioner perceptions align with academic research on QA4AI?
Key findings
- Correctness is the most critical QA4AI property, followed by model relevance, efficiency, and deployability.
- Practitioners report significant trade-offs between latency, cost, and accuracy, with model compression as a key solution.
- Fairness, security, and transferability are less prioritized compared to correctness and efficiency.
- Ten QA4AI practices were well-supported by practitioners, including version control for data and models and automated testing of model inputs.
- Eight practices were marginally agreed upon, such as logging model predictions and using A/B testing for model comparison.
- Practitioners acknowledged academic tools and techniques but noted gaps in real-world adoption due to complexity and integration challenges.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.