Skip to main content
QUICK REVIEW

[论文解读] Explainability in reinforcement learning: perspective and position

Agneza Krajna, Mario Brčić|arXiv (Cornell University)|Mar 22, 2022
Explainable Artificial Intelligence (XAI)被引用 16
一句话总结

本文提出了一种新颖的统一分类法与三支柱框架——主动性、风险态度和认识论约束——用于可解释强化学习(XRL),以解决强化学习中系统性解释方法的缺失问题。该框架在最短路径变体问题上得到验证,强调策略层面的解释而非单一动作,以提升在安全关键应用中的信任度与透明度。

ABSTRACT

Artificial intelligence (AI) has been embedded into many aspects of people's daily lives and it has become normal for people to have AI make decisions for them. Reinforcement learning (RL) models increase the space of solvable problems with respect to other machine learning paradigms. Some of the most interesting applications are in situations with non-differentiable expected reward function, operating in unknown or underdefined environment, as well as for algorithmic discovery that surpasses performance of any teacher, whereby agent learns from experimental experience through simple feedback. The range of applications and their social impact is vast, just to name a few: genomics, game-playing (chess, Go, etc.), general optimization, financial investment, governmental policies, self-driving cars, recommendation systems, etc. It is therefore essential to improve the trust and transparency of RL-based systems through explanations. Most articles dealing with explainability in artificial intelligence provide methods that concern supervised learning and there are very few articles dealing with this in the area of RL. The reasons for this are the credit assignment problem, delayed rewards, and the inability to assume that data is independently and identically distributed (i.i.d.). This position paper attempts to give a systematic overview of existing methods in the explainable RL area and propose a novel unified taxonomy, building and expanding on the existing ones. The position section describes pragmatic aspects of how explainability can be observed. The gap between the parties receiving and generating the explanation is especially emphasized. To reduce the gap and achieve honesty and truthfulness of explanations, we set up three pillars: proactivity, risk attitudes, and epistemological constraints. To this end, we illustrate our proposal on simple variants of the shortest path problem.

研究动机与目标

  • 弥补可解释强化学习(XRL)与监督学习中可解释AI相比的关键差距。
  • 克服强化学习可解释性中的挑战,如延迟奖励、信用分配与非独立同分布数据。
  • 通过引入原则性、以用户为中心的框架,缩小解释生成者与接收者之间的差距。
  • 为安全关键领域中的RL系统建立真实、可靠且可操作的解释标准。
  • 为在多样化应用场景与用户群体中评估XRL方法提供概念基础。

提出的方法

  • 提出一种新颖的四维分类法用于XRL:时间范围(反应式/主动式)、环境类型(确定性/随机性)、策略类型(确定性/随机性)与智能体数量。
  • 提出三大支柱以实现真实解释:主动性(前瞻性策略解释)、风险态度(个性化风险敏感性)与认识论约束(决策者计算能力的限制)。
  • 将该框架应用于简单的最短路径问题变体,通过改变环境、策略与智能体类型来说明解释设计。
  • 采用结构因果模型、奖励分解、层次化策略与关系强化学习作为核心解释技术。
  • 强调策略层面的解释而非动作层面,以支持长期信任与系统理解。
  • 整合认识论约束以反映人类与机器推理中的实际限制,避免过度乐观或误导性解释。

实验结果

研究问题

  • RQ1如何在时间与范围之外,系统性地对强化学习中的可解释性进行分类?
  • RQ2由于延迟奖励与信用分配,生成真实、可靠且与用户相关解释的关键挑战是什么?
  • RQ3解释如何考虑决策者的个体风险偏好与计算约束?
  • RQ4在近似强化学习算法中,智能体与问题之间的本体论与认识论鸿沟如何破坏解释的有效性?
  • RQ5与反应式、聚焦于动作的解释相比,主动式解释在提升用户信任与系统透明度方面有何优势?

主要发现

  • 所提出的三支柱框架——主动性、风险态度与认识论约束——为真实可信的XRL解释提供了原则性基础。
  • 用户更偏好对策略的主动解释,而非反应式、聚焦于特定动作的解释,因为前者有助于长期理解与信任建立。
  • 当前的XRL方法往往无法解释策略,而仅关注单个动作,这限制了透明度与可用性。
  • 忽略风险态度或计算约束的解释可能导致误导性或不切实际的建议,尤其在安全关键领域中。
  • 该框架揭示,即使理论上信息可用,实际计算限制也可能要求解释对缺失知识保持敏感。
  • 本文指出,亟需开发新型评估框架(如可解释性说明表)用于XRL方法,类似于监督学习中XAI所采用的框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。