Skip to main content
QUICK REVIEW

[논문 리뷰] Machine Explanations and Human Understanding

Chacha Chen, Feng Shi|arXiv (Cornell University)|2022. 02. 08.
Explainable Artificial Intelligence (XAI)인용 수 6
한 줄 요약

이 논문은 인간-AI 의사결정에서 기계적 설명과 인간 이해를 연결짓는 공식적인 이론적 프레임워크를 제안하며, 설명이 인간의 작업 결정이나 오류 이해를 향상시킬 수 있는 것은 임의의 작업 특화 인간 직관에 기반할 때에만 가능하다는 것을 입증한다. 핵심 기여는 효과적인 인간-AI 협업을 위해서는 인간의 직관을 명시적으로 모델링할 필요가 있으며, 이는 통제된 인간 대상 실험을 통해 검증되었으며, 이러한 직관이 없을 경우 AI에 대한 과도한 의존이 발생함을 보여준다.

ABSTRACT

Explanations are hypothesized to improve human understanding of machine learning models and achieve a variety of desirable outcomes, ranging from model debugging to enhancing human decision making. However, empirical studies have found mixed and even negative results. An open question, therefore, is under what conditions explanations can improve human understanding and in what way. Using adapted causal diagrams, we provide a formal characterization of the interplay between machine explanations and human understanding, and show how human intuitions play a central role in enabling human understanding. Specifically, we identify three core concepts of interest that cover all existing quantitative measures of understanding in the context of human-AI decision making: task decision boundary, model decision boundary, and model error. Our key result is that without assumptions about task-specific intuitions, explanations may potentially improve human understanding of model decision boundary, but they cannot improve human understanding of task decision boundary or model error. To achieve complementary human-AI performance, we articulate possible ways on how explanations need to work with human intuitions. For instance, human intuitions about the relevance of features (e.g., education is more important than age in predicting a person's income) can be critical in detecting model error. We validate the importance of human intuitions in shaping the outcome of machine explanations with empirical human-subject studies. Overall, our work provides a general framework along with actionable implications for future algorithmic development and empirical experiments of machine explanations.

연구 동기 및 목표

  • 기계적 설명이 인간 이해를 향상시키는 조건을 규명함으로써, 서로 배치된 경험적 결과를 해결하기 위해.
  • 세 핵심 이해 개념인 작업 결정 경계, 모델 결정 경계, 모델 오류 간의 차이를 체계화하기 위해.
  • 인간의 직관에 대한 가정이 없이선 기계적 설명이 작업 수준의 결정이나 오류 경계에 대한 인간 이해를 향상시킬 수 없다는 것을 입증하기 위해.
  • 통제된 인간 대상 실험을 통해 이론적 주장의 타당성을 검증하기 위해.
  • 연구 설계 및 설명 시스템 개발 과정에서 인간의 직관을 명시적으로 표현할 것을 주장하기 위해.

제안 방법

  • 인간이 작업/모델 경계를 근사하는 방식과 기계적 설명 간의 관계를 모델링하기 위해 적응된 원인도를 개발한다.
  • 개입(예: 모델 예측 공개)이 인간의 이해를 어떻게 형성하는지 표현하기 위해 공식적인 $\text{show}$ 연산자를 도입한다.
  • 작업 맥락에 따라 이해를 두 유형으로 분류한다: 모델 이해(모델에 대한 이해)와 탐색(작업에 대한 이해).
  • 설명이 인간의 직관(예: 특징 중요도 또는 상관관계 방향 등)과 일치해야만 작업 수준의 결정 이해를 향상시킬 수 있다는 제안을 한다.
  • 인간의 직관을 고립하고 통제할 수 있도록 워즈더오즈 실험 설계를 활용하며, 직감이 있는 그룹과 없는 그룹 간의 일치율을 비교한다.
  • 독립표본 t-검정과 대응표본 t-검정을 사용하여, 다양한 조건 하에서 인간-AI 결정 간 일치율을 비교하고, 설명의 일관성과 일치도를 측정한다.
Figure 1 : Illustration of the three core concepts using a binary classification problem. Task decision boundary (dashed line) defines the ground-truth mapping from inputs to labels. Model decision boundary (solid line) defines the model predictions. Model error (highlighted area) represents where t
Figure 1 : Illustration of the three core concepts using a binary classification problem. Task decision boundary (dashed line) defines the ground-truth mapping from inputs to labels. Model decision boundary (solid line) defines the model predictions. Model error (highlighted area) represents where t

실험 결과

연구 질문

  • RQ1기계적 설명이 기계 학습 모델에 대한 인간 이해를 향상시킬 수 있는 조건은 무엇인가?
  • RQ2이론적 기대에 비해 경험적 연구에서 설명에 대해 혼합되거나 부정적인 결과가 보고되는 이유는 무엇인가?
  • RQ3작업 특화 인간의 직관은 인간-AI 의사결정에서 기계적 설명의 효과성에 어떻게 영향을 미치는가?
  • RQ4인간의 직관에 대한 가정 없이도 설명이 작업 결정 경계나 모델 오류에 대한 이해를 향상시킬 수 있는가?
  • RQ5인간의 직관이 부재할 경우 AI 예측에 대한 과도한 의존은 어느 정도까지 발생하는가?

주요 결과

  • 인간의 직관에 대한 가정이 없을 경우, 기계적 설명은 작업 결정 경계나 모델 오류에 대한 이해 향상에 기여할 수 없으며, 오직 모델 결정 경계에 대한 이해 향상에만 기여한다.
  • 익명 그룹(직관 없음)에서는 참가자들이 AI 예측과 일치하는 비율이 70.64%였고, 이는 직관이 있는 일반 그룹(54.32%)보다 유의미하게 높아, 직관이 없을 경우 AI에 대한 과도한 의존이 발생함을 시사한다.
  • 직관이 존재할 경우, 일관된 쌍(AB)에 대해 설명과의 일치율은 90.71%였고, 불일치하는 쌍(CD)에 대해서는 단지 25.00%였으며, 이는 올바른 직관과의 강한 일치를 보여준다.
  • 인간의 직관과 일치할 경우 설명의 일관성이 유의미하게 높아졌다: AB 그룹은 90.71%의 일치율을 기록했고, EF 그룹은 22.14%였으며, p값 < 0.001이었다.
  • t-검정 결과, 다양한 조건 간 일치율의 차이가 통계적으로 유의미하게 나타났다(p < 0.001), 이는 직관이 설명 효과성에 기여한다는 것을 검증한다.
  • 이 연구는 이론적 프레임워크가 예측한 바와 같이, 작업 특화 인간의 직관이 없이선 인간+AI의 보완적 성능(인간+AI > 인간 및 > AI)은 불가능하다는 것을 확인한다.
Figure 2 : Visualizing the relations between local variables, core functions, and human approximations of them.
Figure 2 : Visualizing the relations between local variables, core functions, and human approximations of them.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.