[논문 리뷰] Trust in AI: Interpretability is not necessary or sufficient, while black-box interaction is necessary and sufficient
이 논문은 해석 가능성(interpretability)이 인간의 AI 신뢰에 필수적이지도 않고 충분히 신뢰할 만한 AI 시스템을 구축하는 데 필수적인 요소가 아니라고 주장한다. 대신 블랙박스 상호작용(black-box interaction)—특히 모델를 원할 때마다 실행하고 테스트할 수 있는 능력—이 신뢰에 대해 필수적이며 충분한 조건임을 밝힌다. 저자들은 모델의 분포 외(out-of-distribution) 및 임무 외(out-of-task) 성능에 대한 경험적이고 이론적인 증거를 종합하여 신뢰성 평가를 가능하게 하는 행동 인증서(behavior certificate) 프레임워크를 제안하며, 모델 내부 구조를 이해하는 데서 모델 행동을 이해하는 데로 초점을 이동시킨다.
The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health, and humanity, yet confusion surrounds foundational concepts. What does it mean to trust an AI, and how do humans assess AI trustworthiness? What are the mechanisms for building trustworthy AI? And what is the role of interpretable ML in trust? Here, we draw from statistical learning theory and sociological lenses on human-automation trust to motivate an AI-as-tool framework, which distinguishes human-AI trust from human-AI-human trust. Evaluating an AI's contractual trustworthiness involves predicting future model behavior using behavior certificates (BCs) that aggregate behavioral evidence from diverse sources including empirical out-of-distribution and out-of-task evaluation and theoretical proofs linking model architecture to behavior. We clarify the role of interpretability in trust with a ladder of model access. Interpretability (level 3) is not necessary or even sufficient for trust, while the ability to run a black-box model at-will (level 2) is necessary and sufficient. While interpretability can offer benefits for trust, it can also incur costs. We clarify ways interpretability can contribute to trust, while questioning the perceived centrality of interpretability to trust in popular discourse. How can we empower people with tools to evaluate trust? Instead of trying to understand how a model works, we argue for understanding how a model behaves. Instead of opening up black boxes, we should create more behavior certificates that are more correct, relevant, and understandable. We discuss how to build trusted and trustworthy AI responsibly.
연구 동기 및 목표
- 해석 가능성의 인간이 AI 시스템을 신뢰하는 데서 차지하는 기초적 역할을 명확히 하기.
- 신뢰할 만한 AI를 구축하기 위해 해석 가능성은 필수적이라는 널리 퍼진 가정에 도전하기.
- 행동 인증서를 통한 모델 행동 평가로 해석 가능성에서의 전환을 제안하기.
- 블랙박스 상호작용을 AI에서 신뢰를 평가하고 가능하게 하는 핵심 메커니즘으로 설정하기.
- 신뢰 최적화보다는 계약 인식 모델 설계, 강건성 테스팅, 신뢰 조정을 지지하기.
제안 방법
- 인간-AI 신뢰와 인간-AI-인간 신뢰를 구분하는 AI-as-tool 프레임워크 도입.
- 행동 인증서(BCs) 제안 — 경험적 분포 외 및 임무 외 평가를 종합하고 아키텍처와 행동 간 이론적 증명을 연결.
- 모델 액세스 계층 정의: 수준 2(블랙박스 상호작용)는 신뢰에 필수적이고 충분함; 수준 3(해석 가능성)는 필수적이지도 않고 충분하지도 않음.
- 민감도 분석, 적대적 테스팅, 하이퍼파라미터 최적화와 같은 블랙박스 액세스만을 사용한 모델 디버깅 및 과학적 발견을 지지.
- 해석 가능성 중심 접근 방식을 더 정확하고 관련성 있으며 이해하기 쉬운 행동 인증서로 대체할 것을 권장.
- 신뢰 최적화보다는 모델 카드, 공급자 준수 선언, 소프트웨어 공학에서 유래한 강건성 테스팅 원칙을 활용한 신뢰 조정을 촉진.
실험 결과
연구 질문
- RQ1인간의 AI 시스템에 대한 신뢰에 대해 진정으로 필수적이고 충분한 메커니즘은 무엇인가?
- RQ2해석 가능성과 블랙박스 상호작용 중 어느 것이 신뢰를 더 잘 가능하게 하는가?
- RQ3내부 모델 메커니즘을 이해하지 않더라도 모델의 신뢰성을 평가할 수 있는가?
- RQ4행동 인증서는 신뢰할 만한 AI를 구축하는 데 어떤 역할을 하는가?
- RQ5해석 가능성 없이 과학적 발견과 모델 디버깅은 어떻게 진행될 수 있는가?
주요 결과
- 해석 가능성은 널리 중시되지만, 인간의 AI 신뢰에 대해 필수적이지도 않고 충분하지도 않다.
- 블랙박스 상호작용 — 특히 원할 때마다 모델를 실행하고 테스트할 수 있는 능력 — 이 신뢰에 대해 필수적이며 충분하다.
- 모델 행동에 대한 경험적 및 이론적 증거를 종합한 행동 인증서는 해석 가능성보다 더 효과적인 신뢰 지표이다.
- 민감도 분석 및 하이퍼파라미터 최적화와 같은 블랙박스 액세스만으로도 모델 디버깅 및 과학적 발견을 효과적으로 수행할 수 있다.
- 특히 단순화 학습에 취약한 모델에서는 해석 가능성으로 잘못된 통찰을 유도해 신뢰를 저해할 수 있다.
- 실제 응용에서 해석 가능성 방법보다는 강건성 테스팅, 보류된 데이터 평가, 모델 카드가 더 큰 영향을 미친다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.