Skip to main content
QUICK REVIEW

[論文レビュー] The MineRL BASALT Competition on Learning from Human Feedback

Rohin Shah, Cody Wild|arXiv (Cornell University)|Jul 5, 2021
Reinforcement Learning in Robotics参考文献 31被引用数 4
ひとこと要約

MineRL BASALT コンペティションは、事前に定義された報酬関数の代わりに人間のフィードバックを用いてAIエージェントを訓練するためのベンチマークを提唱し、Minecraftにおける複雑で自然言語で記述されたタスクに注力している。タスク完了の観点から人間の判断による評価がなされ、価値の整合性やオープンエンドな環境におけるスケーラブルな模倣学習および好み学習の分野における研究を前進させる。

ABSTRACT

The last decade has seen a significant increase of interest in deep learning research, with many public successes that have demonstrated its potential. As such, these systems are now being incorporated into commercial products. With this comes an additional challenge: how can we build AI systems that solve tasks where there is not a crisp, well-defined specification? While multiple solutions have been proposed, in this competition we focus on one in particular: learning from human feedback. Rather than training AI systems using a predefined reward function or using a labeled dataset with a predefined set of categories, we instead train the AI system using a learning signal derived from some form of human feedback, which can evolve over time as the understanding of the task changes, or as the capabilities of the AI system improve. The MineRL BASALT competition aims to spur forward research on this important class of techniques. We design a suite of four tasks in Minecraft for which we expect it will be hard to write down hardcoded reward functions. These tasks are defined by a paragraph of natural language: for example, "create a waterfall and take a scenic picture of it", with additional clarifying details. Participants must train a separate agent for each task, using any method they want. Agents are then evaluated by humans who have read the task description. To help participants get started, we provide a dataset of human demonstrations on each of the four tasks, as well as an imitation learning baseline that leverages these demonstrations. Our hope is that this competition will improve our ability to build AI systems that do what their designers intend them to do, even when the intent cannot be easily formalized. Besides allowing AI to solve more tasks, this can also enable more effective regulation of AI systems, as well as making progress on the value alignment problem.

研究の動機と目的

  • 手動で作成された報酬関数の代替手段として、人間からのフィードバックによる学習(LfHF)の研究を前進させること。
  • 仕様が曖昧または不完全なタスクにおけるAIエージェントの訓練という課題に取り組むこと。
  • 報酬の最大化ではなく、タスク完了の観点から人間の判断に基づいてエージェントを評価するベンチマークを構築すること。
  • 複雑で自然言語で記述された指示を理解し実行できる、人間と整合性のあるスケーラブルなAIシステムの開発を促進すること。
  • 多様なフィードバックモダリティを用いて人間の意図を組み込むことで、報酬のあいまいさや誤った報酬の利用(reward hacking)を低減すること。

提案手法

  • 参加者は、タスク用に事前に定義された報酬関数が提供されない状態で、任意の方法でエージェントを訓練する。
  • タスクは「滝を作成し、その風景を写真に撮る」といった自然言語による記述で定義される。
  • 人間の評価者がタスク記述に基づいてエージェントのパフォーマンスを判断し、期待される行動と整合させる。
  • 模倣学習のベースラインを支援するために、人間のデモンストレーションのデータセットが提供される。
  • コンペティションでは、多様なエージェント行動と目的をサポートする豊富なオープンエンドな世界を備えたMineRL環境が使用される。
  • 評価は人間のレーティング担当者が実施され、環境設計に起因するバイアスを低減する。

実験結果

リサーチクエスチョン

  • RQ1人間からのフィードバックを用いて訓練されたエージェントは、オープンエンドな環境における複雑で自然言語で記述されたタスクにどの程度一般化できるか?
  • RQ2報酬関数が曖昧または誤解を招く状況では、人間からのフィードバックによる学習が、従来の報酬設計を上回る性能を発揮できるか?
  • RQ3模倣学習と好みベースの手法は、現実に近い複雑なタスクにおいて、人間の意図との整合性をどの程度向上できるか?
  • RQ4自動報酬信号と比較して、人間によるエージェント行動の評価は、タスク成功の測定においてどの程度有効か?
  • RQ5オープンワールド環境において、最も強固で整合性の高いエージェント行動をもたらすフィードバックモダリティ(例:デモンストレーション、比較)は何か?

主な発見

  • コンペティションは、事前に定義された報酬関数なしに、自然言語で記述された複雑なMinecraftタスクを人間のフィードバックによってエージェントが遂行できることを明確に示した。
  • 提供された人間のデモンストレーションから出発する模倣学習のベースラインを用いて訓練されたエージェントは中程度のパフォーマンスを示し、さらなる改善の強力な出発点であることがわかった。
  • 人間による評価から、エージェントがタスク記述の微細な側面を理解できていないケースが多く見られ、より良い整合性技術の必要性が浮き彫りになった。
  • Minecraftのオープンエンドな性質が、報酬関数が使用された場合に意図しない行動が報酬として与えられがちなことから、LfHF手法の評価に適したテストベッドであった。
  • 専用のDiscordフォーラムを通じて、参加者間の協働的で活発な研究コミュニティが形成され、参加者の関与度と知識共有が促進された。
  • 人間による評価を主たる指標として採用することで、環境設計に起因するバイアスが低減され、現実世界のタスク成功をより的確に反映するようになった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。