[논문 리뷰] Quality Assurance for Artificial Intelligence: A Study of Industrial Concerns, Challenges and Best Practices
이 연구는 15명의 전문가 인터뷰와 50명의 설문 조사로 산업 현장에서의 인공지능 품질보증(QA4AI) 관련 고려사항, 과제, 최선의 실천 방안을 조사한다. 정확성이 최우선으로 여겨지며, 그 다음으로 모델의 관련성, 효율성, 구현 가능성 순이다. 연구는 QA4AI 실천 방안 21가지를 제안하며, 이 중 10개는 강력한 지지, 8개는 약간의 공감대를 확보하였다. 이는 산업 현장 적용을 위한 실용적인 체크리스트를 제공한다.
Quality Assurance (QA) aims to prevent mistakes and defects in manufactured products and avoid problems when delivering products or services to customers. QA for AI systems, however, poses particular challenges, given their data-driven and non-deterministic nature as well as more complex architectures and algorithms. While there is growing empirical evidence about practices of machine learning in industrial contexts, little is known about the challenges and best practices of quality assurance for AI systems (QA4AI). In this paper, we report on a mixed-method study of QA4AI in industry practice from various countries and companies. Through interviews with fifteen industry practitioners and a validation survey with 50 practitioner responses, we studied the concerns as well as challenges and best practices in ensuring the QA4AI properties reported in the literature, such as correctness, fairness, interpretability and others. Our findings suggest correctness as the most important property, followed by model relevance, efficiency and deployability. In contrast, transferability (applying knowledge learned in one task to another task), security and fairness are not paid much attention by practitioners compared to other properties. Challenges and solutions are identified for each QA4AI property. For example, interviewees highlighted the trade-off challenge among latency, cost and accuracy for efficiency (latency and cost are parts of efficiency concern). Solutions like model compression are proposed. We identified 21 QA4AI practices across each stage of AI development, with 10 practices being well recognized and another 8 practices being marginally agreed by the survey practitioners.
연구 동기 및 목표
- 정확성, 공정성, 해석 가능성과 같은 QA4AI 성질에 대한 산업 현장의 인식을 이해하기 위해.
- 실제 AI 개발 현장에서 각 QA4AI 성질을 확보하기 위한 주요 과제와 해결책을 특정하기 위해.
- AI 개발 라이프사이클 전반에 걸쳐 QA4AI의 최선의 실천 방안을 도출하고 검증하기 위해.
- 학술 연구와 산업 응용 간의 격차를 AI 품질보증 분야에서 해소하기 위해.
제안 방법
- 다양한 기업과 국가의 15명의 AI 전문가와 함께 반구조화된 인터뷰를 실시하였다.
- 결과의 공감대를 평가하기 위해 추가로 50명의 산업 현장 전문가에게 설문 조사를 실시하였다.
- QA4AI 성질을 9단계 AI 개발 워크플로우에 매핑하여 포괄적인 커버리지 확보를 위해 노력하였다.
- 인터뷰 데이터와 설문 응답을 바탕으로 주제 분석을 통해 QA4AI 실천 방안 21가지를 도출하였다.
- 설문 공감도 기반으로 실천 방안을 '강력히 지지됨'(10개) 또는 '약간의 공감' (8개)로 분류하였다.
- 전문가들이 보고한 트레이드오프(예: 지연 시간 대 비용 대 정확도)와 도구 사용 사례를 분석하였다.
실험 결과
연구 질문
- RQ1산업 현장의 전문가들은 QA4AI 성질의 중요도를 어떻게 순위 매기나?
- RQ2실제 현장에서 각 QA4AI 성질을 확보하기 위한 주요 과제와 해결책은 무엇인가?
- RQ3AI 개발 라이프사이클 전반에 걸쳐 인정되고 도입된 최선의 실천 방안은 무엇인가?
- RQ4전문가들의 인식은 QA4AI에 관한 학술 연구와 얼마나 일치하는가?
주요 결과
- 정확성이 가장 중요한 QA4AI 성질이며, 그 다음으로 모델의 관련성, 효율성, 구현 가능성이다.
- 전문가들은 지연 시간, 비용, 정확도 사이에 상당한 트레이드오프를 경험하며, 모델 압축이 핵심 해결책으로 작용한다.
- 정확성과 효율성에 비해 공정성, 보안성, 이식 가능성은 덜 우선시된다.
- 10개의 QA4AI 실천 방안이 전문가들 사이에서 강력히 지지되었으며, 데이터 및 모델의 버전 관리와 모델 입력의 자동 테스트가 포함되어 있다.
- 8개의 실천 방안은 약간의 공감대를 확보하였으며, 모델 예측 로깅과 모델 비교를 위한 A/B 테스트 사용 등이 포함되어 있다.
- 전문가들은 학술 분야의 도구와 기법을 인정하지만, 복잡성과 통합 과제로 인해 실제 현장 적용에서 격차가 존재한다고 지적하였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.