Skip to main content
QUICK REVIEW

[논문 리뷰] Explainability in reinforcement learning: perspective and position

Agneza Krajna, Mario Brčić|arXiv (Cornell University)|2022. 03. 22.
Explainable Artificial Intelligence (XAI)인용 수 16
한 줄 요약

이 논문은 강화학습(Reinforcement Learning, RL)에서 체계적인 설명 방법의 부족을 해결하기 위해 새로운 통합 분류 체계와 세 축심의 프레임워크—사전성, 위험 태도, 인지론적 제약 조건—을 제안한다. 이는 설명 가능한 강화학습(XRL)을 위한 것이다. 프레임워크는 단순한 최단경로 문제 변형에 적용되어, 단일 행동이 아닌 정책 수준의 설명을 강조함으로써 안전 중심 응용 분야에서 신뢰성과 투명성을 향상시킨다.

ABSTRACT

Artificial intelligence (AI) has been embedded into many aspects of people's daily lives and it has become normal for people to have AI make decisions for them. Reinforcement learning (RL) models increase the space of solvable problems with respect to other machine learning paradigms. Some of the most interesting applications are in situations with non-differentiable expected reward function, operating in unknown or underdefined environment, as well as for algorithmic discovery that surpasses performance of any teacher, whereby agent learns from experimental experience through simple feedback. The range of applications and their social impact is vast, just to name a few: genomics, game-playing (chess, Go, etc.), general optimization, financial investment, governmental policies, self-driving cars, recommendation systems, etc. It is therefore essential to improve the trust and transparency of RL-based systems through explanations. Most articles dealing with explainability in artificial intelligence provide methods that concern supervised learning and there are very few articles dealing with this in the area of RL. The reasons for this are the credit assignment problem, delayed rewards, and the inability to assume that data is independently and identically distributed (i.i.d.). This position paper attempts to give a systematic overview of existing methods in the explainable RL area and propose a novel unified taxonomy, building and expanding on the existing ones. The position section describes pragmatic aspects of how explainability can be observed. The gap between the parties receiving and generating the explanation is especially emphasized. To reduce the gap and achieve honesty and truthfulness of explanations, we set up three pillars: proactivity, risk attitudes, and epistemological constraints. To this end, we illustrate our proposal on simple variants of the shortest path problem.

연구 동기 및 목표

  • 설문 학습에서의 설명 가능성 AI에 비해 설명 가능한 강화학습(XRL)의 핵심적 격차를 해결한다.
  • 강화학습의 설명 가능성에서 지연된 보상, 책임 할당, 비i.i.d. 데이터 등의 과제를 극복한다.
  • 설명 생성자와 수신자 간 격차를 줄이기 위해 원칙적이고 사용자 중심의 프레임워크를 도입한다.
  • 안전 중심 영역에서 진정성 있고 신뢰할 수 있으며 실행 가능한 설명을 위한 표준을 수립한다.
  • 다양한 응용 시나리오와 사용자 프로파일에 걸쳐 XRL 방법을 평가하기 위한 개념적 기반을 제공한다.

제안 방법

  • XRL에 대한 새로운 사원축 분류 체계 도입: 시간 범위(반응형/사전성), 환경 유형(결정론적/스토케스틱), 정책 유형(결정론적/스토케스틱), 에이전트 수량.
  • 진정성 있는 설명을 위한 세 축심 도입: 사전성(예측 가능한 정책 설명), 위험 태도(개인화된 위험 민감도), 지식론적 제약 조건(결정자의 계산적 한계).
  • 환경, 정책, 에이전트 유형을 변화시켜 단순한 최단경로 문제 변형에 프레임워크를 적용하여 설명 설계의 예시를 제시한다.
  • 구조적 인과 모델, 보상 분해, 계층적 정책, 관계 기반 강화학습을 핵심 설명 기법으로 사용한다.
  • 장기적인 신뢰성과 시스템 이해를 지원하기 위해 행동 수준 설명보다 정책 수준 설명을 강조한다.
  • 사용자와 기계의 추론에서 현실적인 제약 조건을 반영하기 위해 지식론적 제약 조건을 통합하여 과도하게 낙관적이거나 오해의 소지가 있는 설명을 방지한다.

실험 결과

연구 질문

  • RQ1시간과 범위를 넘어서 설명 가능성의 체계적 분류는 어떻게 가능할 수 있는가?
  • RQ2지연된 보상과 책임 할당으로 인해 강화학습에서 진정성 있고 신뢰할 수 있으며 사용자 관련 설명을 생성하는 데 있어 핵심 과제는 무엇인가?
  • RQ3설명은 결정자의 개인적 위험 선호도와 계산적 제약 조건을 어떻게 반영할 수 있는가?
  • RQ4근사 강화학습 알고리즘에서 에이전트와 문제 사이의 철학적 및 지식론적 격차는 설명의 타당성을 어떻게 약화시키는가?
  • RQ5반응형 행동 중심 설명에 비해 사전성 설명은 사용자 신뢰와 시스템 투명성을 어떻게 향상시킬 수 있는가?

주요 결과

  • 사전성, 위험 태도, 지식론적 제약 조건으로 구성된 제안된 세 축심 프레임워크는 신뢰할 수 있고 진정성 있는 XRL 설명을 위한 원칙적인 기반을 제공한다.
  • 사용자들은 반응형 행동 중심 설명보다 사전성 정책 설명을 선호하며, 이는 장기적인 이해와 신뢰를 지원하기 때문이다.
  • 현재 XRL 방법은 일반적으로 정책을 설명하지 못하고 단일 행동에 집중함으로써 투명성과 사용성에 한계를 가진다.
  • 위험 태도나 계산적 제약 조건을 忽시한 설명은 특히 안전 중심 분야에서 오해의 소지가 있거나 실현 가능성이 없는 권고로 이어질 수 있다.
  • 정보가 이론적으로는 가능하더라도 계산의 실질적 제약으로 인해 일부 지식이 생략된 경우 설명이 이를 반영해야 할 필요가 있음을 프레임워크가 드러낸다.
  • 이 논문은 설문 학습에서 XAI에 사용되는 것과 유사한 '설명 가능성 팩트 시트'와 같은 새로운 평가 프레임워크의 필요성을 제기한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.