[论文解读] Quality Assurance for Artificial Intelligence: A Study of Industrial Concerns, Challenges and Best Practices
本研究通过与15位从业者访谈及对50位从业者的调查,探讨了工业界在人工智能质量保障(QA4AI)方面的关注点、挑战与最佳实践。研究发现,正确性是首要关注点,其次为模型相关性、效率与可部署性。研究提出了21项QA4AI实践,其中10项获得充分支持,8项达成部分共识,为工业界实施提供了实用检查清单。
Quality Assurance (QA) aims to prevent mistakes and defects in manufactured products and avoid problems when delivering products or services to customers. QA for AI systems, however, poses particular challenges, given their data-driven and non-deterministic nature as well as more complex architectures and algorithms. While there is growing empirical evidence about practices of machine learning in industrial contexts, little is known about the challenges and best practices of quality assurance for AI systems (QA4AI). In this paper, we report on a mixed-method study of QA4AI in industry practice from various countries and companies. Through interviews with fifteen industry practitioners and a validation survey with 50 practitioner responses, we studied the concerns as well as challenges and best practices in ensuring the QA4AI properties reported in the literature, such as correctness, fairness, interpretability and others. Our findings suggest correctness as the most important property, followed by model relevance, efficiency and deployability. In contrast, transferability (applying knowledge learned in one task to another task), security and fairness are not paid much attention by practitioners compared to other properties. Challenges and solutions are identified for each QA4AI property. For example, interviewees highlighted the trade-off challenge among latency, cost and accuracy for efficiency (latency and cost are parts of efficiency concern). Solutions like model compression are proposed. We identified 21 QA4AI practices across each stage of AI development, with 10 practices being well recognized and another 8 practices being marginally agreed by the survey practitioners.
研究动机与目标
- 理解工业界对QA4AI属性(如正确性、公平性与可解释性)的看法。
- 识别在实际AI开发中确保各项QA4AI属性的关键挑战与解决方案。
- 提取并验证贯穿AI开发全生命周期的QA4AI最佳实践。
- 弥合学术研究与工业应用在AI质量保障之间的差距。
提出的方法
- 对来自不同公司与国家的15位AI从业者进行了半结构化访谈。
- 向另外50位行业从业者发放验证问卷,以评估其对研究发现的共识程度。
- 将QA4AI属性映射至9阶段的AI开发工作流,以确保全面覆盖。
- 通过访谈数据与调查反馈的主题分析,识别出21项QA4AI实践。
- 根据调查共识,将实践分类为“充分支持”(10项)或“部分共识”(8项)。
- 分析从业者报告的权衡关系(如延迟 vs. 成本 vs. 准确率)及工具使用情况。
实验结果
研究问题
- RQ1行业从业者如何对不同QA4AI属性的重要性进行排序?
- RQ2在实践中,确保各项QA4AI属性的关键挑战与解决方案是什么?
- RQ3在AI开发全生命周期中,哪些最佳实践被广泛认可并采纳?
- RQ4从业者观点与学术界关于QA4AI的研究是否一致?
主要发现
- 正确性是最重要的QA4AI属性,其次为模型相关性、效率与可部署性。
- 从业者报告在延迟、成本与准确率之间存在显著权衡,模型压缩是关键解决方案。
- 与正确性和效率相比,公平性、安全性和可迁移性优先级较低。
- 10项QA4AI实践获得从业者充分支持,包括数据与模型的版本控制,以及模型输入的自动化测试。
- 8项实践达成部分共识,如记录模型预测结果,以及使用A/B测试进行模型对比。
- 从业者承认学术界工具与技术的存在,但指出由于复杂性与集成挑战,实际应用中存在差距。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。