Skip to main content
QUICK REVIEW

[論文レビュー] A Framework for Decision-Theoretic Planning I: Combining the Situation Calculus, Conditional Plans, Probability and Utility

David Poole|arXiv (Cornell University)|Feb 13, 2013
Logic, Reasoning, and Knowledge参考文献 27被引用数 11
ひとこと要約

本稿は、論理的行動表現、条件付き計画、確率、効用を統合する包括的な枠組みを提示する。確率的フレーム問題と独立選択論理を用いて不確実性をモデル化し、期待効用を計算することで、確率的STRIPSと比較して指数的空間効率を達成した不確実性下での最適計画を可能にする。

ABSTRACT

This paper shows how we can combine logical representations of actions and decision theory in such a manner that seems natural for both. In particular we assume an axiomatization of the domain in terms of situation calculus, using what is essentially Reiter's solution to the frame problem, in terms of the completion of the axioms defining the state change. Uncertainty is handled in terms of the independent choice logic, which allows for independent choices and a logic program that gives the consequences of the choices. As part of the consequences are a specification of the utility of (final) states. The robot adopts robot plans, similar to the GOLOG programming language. Within this logic, we can define the expected utility of a conditional plan, based on the axiomatization of the actions, the uncertainty and the utility. The ?planning' problem is to find the plan with the highest expected utility. This is related to recent structured representations for POMDPs; here we use stochastic situation calculus rules to specify the state transition function and the reward/value function. Finally we show that with stochastic frame axioms, actions representations in probabilistic STRIPS are exponentially larger than using the representation proposed here.

研究の動機と目的

  • 状況計算における論理的行動表現を意思決定理論と統合し、不確実性下での計画を可能にする。
  • 独立選択論理と確率的フレーム問題を用いて不確実性をモデル化し、確率的状態遷移を可能にする。
  • 論理的枠組み内で条件付き計画の期待効用を定義し、最適計画を実現する。
  • 提案された表現が確率的STRIPSと比較して指数的にコンパクトであることを示す。
  • 論理的および確率的推論を用いて意思決定理論的計画の形式的基盤を提供する。

提案手法

  • Reイターのフレーム問題の解決法を用い、状況計算における行動公理の完成を実現する。
  • 独立選択論理を用いて確率的選択とその結果をモデル化する。
  • 確率的状況計算ルールを用いて状態遷移関数と報酬/価値関数を定義する。
  • GOLOGに類似したロボットプログラムとして計画を表現し、条件付き実行をサポートする。
  • 行動公理、不確実性、および効用関数に基づいて計画の期待効用を計算する。
  • 論理的推論を用いて期待効用が最大となる計画を特定する。

実験結果

リサーチクエスチョン

  • RQ1状況計算における論理的行動表現をどのように意思決定理論的計画に拡張できるか。
  • RQ2論理的枠組み内で確率的状態遷移と報酬を最もコンパクトかつ自然に表現する方法は何か。
  • RQ3論理的かつ確率的システム内で条件付き計画の期待効用をどのように評価できるか。
  • RQ4本フレームワークは、確率的STRIPSなどの既存手法と比較して、どのような表現的および計算的利点を有するか。
  • RQ5統合的論理ベースのシステムは、行動、不確実性、および効用を効果的に統合し、最適計画を実現できるか。

主な発見

  • 論理公理と確率的選択を用いることで、条件付き計画の期待効用を定義可能である。
  • 確率的フレーム問題の導入により、確率的STRIPSと比較して状態遷移の表現がよりコンパクトに可能である。
  • 提案手法は、確率的STRIPS表現と比較して指数的空間削減を達成する。
  • 状況計算への効用の統合により、期待効用の最大化による最適計画が可能になる。
  • 確率的整合性を保ちながら、計画に対する構造的論理的推論をサポートする。
  • フレームワークは独立選択論理に基づいて形式的に根拠付けられており、意思決定理論的計画のための状況計算の自然な拡張を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。