Skip to main content
QUICK REVIEW

[論文レビュー] Interpretable-AI Policies using Evolutionary Nonlinear Decision Trees for Discrete Action Systems.

Yashesh Dhebar, Kalyanmoy Deb|arXiv (Cornell University)|Sep 20, 2020
Reinforcement Learning in Robotics参考文献 31被引用数 5
ひとこと要約

本稿では、ブラックボックス型の深層強化学習(DRL)ポリシーを、単純で人間が読みやすいルールに蒸留する解釈可能なAIポリシーを、進化的非線形意思決定木(NLDT)を用いて提案する。二段階最適化手順と再最適化を活用することで、DRLと同等の閉ループ性能を達成しつつ、1ルールあたり1〜4つの非線形項を有するポリシーを生成し、性能を犠牲にすることなく完全な解釈可能性を実現する。

ABSTRACT

Black-box artificial intelligence (AI) induction methods such as deep reinforcement learning (DRL) are increasingly being used to find optimal policies for a given control task. Although policies represented using a black-box AI are capable of efficiently executing the underlying control task and achieving optimal closed-loop performance -- controlling the agent from initial time step until the successful termination of an episode, the developed control rules are often complex and neither interpretable nor explainable. In this paper, we use a recently proposed nonlinear decision-tree (NLDT) approach to find a hierarchical set of control rules in an attempt to maximize the open-loop performance for approximating and explaining the pre-trained black-box DRL (oracle) agent using the labelled state-action dataset. Recent advances in nonlinear optimization approaches using evolutionary computation facilitates finding a hierarchical set of nonlinear control rules as a function of state variables using a computationally fast bilevel optimization procedure at each node of the proposed NLDT. Additionally, we propose a re-optimization procedure for enhancing closed-loop performance of an already derived NLDT. We evaluate our proposed methodologies on four different control problems having two to four discrete actions. In all these problems our proposed approach is able to find simple and interpretable rules involving one to four non-linear terms per rule, while simultaneously achieving on par closed-loop performance when compared to a trained black-box DRL agent. The obtained results are inspiring as they suggest the replacement of complicated black-box DRL policies involving thousands of parameters (making them non-interpretable) with simple interpretable policies. Results are encouraging and motivating to pursue further applications of proposed approach in solving more complex control tasks.

研究の動機と目的

  • ブラックボックス型の深層強化学習(DRL)ポリシーに見られる解釈不能性の問題に対処すること。
  • 事前に訓練されたDRLオラクルエージェントの行動を近似する階層的で解釈可能な非線形制御ルールのセットを構築すること。
  • 再最適化手順を用いて、導出された解釈可能なポリシーの閉ループ性能を向上させること。
  • 複雑さの異なる離散的アクション制御タスクにおいて、本手法を評価し、解釈可能性と競争力のある性能を両立させること。

提案手法

  • 状態変数に基づく非線形意思決定ルールの階層として制御ポリシーを表現する非線形意思決定木(NLDT)アーキテクチャを採用する。
  • 各NLDTノードで、進化的計算と非線形最適化を組み合わせた計算的に効率的な二段階最適化手順を用い、最適な非線形項を学習する。
  • 事前に訓練されたDRLエージェント(オラクルポリシー)から抽出した状態-行動ペアのラベル付きデータセットを用いてNLDTを学習する。
  • 初期のルール抽出後に、NLDTポリシーの閉ループ性能を向上させるために、再最適化手順を適用する。
  • 進化的計算を用いて複雑な非線形意思決定境界を探索しつつ、構造的な木構造ベースのルール表現により解釈可能性を維持する。
  • 各ルールに1〜4つの非線形項に制限することで、モデルの複雑さと性能のバランスを図り、人間が読みやすいようにする。

実験結果

リサーチクエスチョン

  • RQ1進化的最適化を用いて導出された非線形意思決定木は、離散的アクション制御タスクにおいてブラックボックス型DRLポリシーの行動を効果的に近似できるか?
  • RQ2得られたNLDTポリシーは、解釈可能である一方で、元のDRLエージェントと同等の閉ループ性能をどの程度維持できるか?
  • RQ3提案された再最適化手順は、蒸留されたNLDTポリシーの閉ループ性能をどの程度向上させられるか?
  • RQ4本手法は、離散的アクション数が2から4に変化する制御問題に一般化可能であり、解釈可能性と性能を両立できるか?

主な発見

  • 提案されたNLDTベースのポリシーは、評価された4つの制御問題すべてで、元のブラックボックス型DRLエージェントと同等の閉ループ性能を達成した。
  • 蒸留されたポリシーの各制御ルールには、非線形項が1〜4つしか含まれておらず、数千のパラメータを持つDRLポリシーと比較して、顕著に解釈性が向上している。
  • 再最適化手順により、NLDTポリシーの閉ループ性能が有意に向上し、蒸留ポリシーの最適化に有効であることが実証された。
  • 本手法は、複雑なDRLポリシーを、単純で階層的かつ人間が読みやすい意思決定ルールに効果的に蒸留でき、タスクの性能を犠牲にすることなく実現した。
  • 結果から、実世界の制御応用において解釈不能なDRLポリシーに代わる代替手段としての強力な可能性が示された。
  • 本手法は2〜4つの離散的アクションを有する問題においても有効であり、中程度の複雑さの制御タスクへのスケーラビリティを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。