Skip to main content
QUICK REVIEW

[論文レビュー] Cohesion-based Online Actor-Critic Reinforcement Learning for mHealth Intervention

Feiyun Zhu, Peng Liao|arXiv (Cornell University)|Mar 25, 2017
Impact of Technology on Adolescents参考文献 8被引用数 4
ひとこと要約

本稿は、データが乏しい環境における方策学習の安定性を向上させるために、ユーザー間のネットワークコhesivenessを活用する、新しい協調的オンラインアクター・クリティック強化学習フレームワークを提案する。温暖スタート軌道からユーザー類似性ネットワークを学習し、価値関数および方策関数にℓ₂制約を課すことにより、分散型方策学習を著しく改善する。本手法は、複数のmHealthベンチマークで優れた性能を示す。

ABSTRACT

In the wake of the vast population of smart device users worldwide, mobile health (mHealth) technologies are hopeful to generate positive and wide influence on people's health. They are able to provide flexible, affordable and portable health guides to device users. Current online decision-making methods for mHealth assume that the users are completely heterogeneous. They share no information among users and learn a separate policy for each user. However, data for each user is very limited in size to support the separate online learning, leading to unstable policies that contain lots of variances. Besides, we find the truth that a user may be similar with some, but not all, users, and connected users tend to have similar behaviors. In this paper, we propose a network cohesion constrained (actor-critic) Reinforcement Learning (RL) method for mHealth. The goal is to explore how to share information among similar users to better convert the limited user information into sharper learned policies. To the best of our knowledge, this is the first online actor-critic RL for mHealth and first network cohesion constrained (actor-critic) RL method in all applications. The network cohesion is important to derive effective policies. We come up with a novel method to learn the network by using the warm start trajectory, which directly reflects the users' property. The optimization of our model is difficult and very different from the general supervised learning due to the indirect observation of values. As a contribution, we propose two algorithms for the proposed online RLs. Apart from mHealth, the proposed methods can be easily applied or adapted to other health-related tasks. Extensive experiment results on the HeartSteps dataset demonstrates that in a variety of parameter settings, the proposed two methods obtain obvious improvements over the state-of-the-art methods.

研究の動機と目的

  • mHealthにおけるユーザー固有のデータが限られる状況下で、個別方策学習の不安定性を解消すること。
  • ユーザー類似性をネットワークコヒーレンスを通じてモデル化し、類似ユーザー間での情報共有を可能にすること。
  • 事前に定義されたネットワークを仮定せず、データから直接コヒーレンスネットワークを学習するオンラインアクター・クリティックRLフレームワークを構築すること。
  • ネットワーク構造に基づく価値関数および方策関数のℓ₂制約により、バイアスを過剰に増加させることなく、方策学習の分散を低減すること。
  • 実世界のmHealthデータ上での最先端手法と比較して、本手法の優位性を実証すること。

提案手法

  • 類似行動を示すユーザー間で情報を共有する、ネットワークコヒーレンス制約付きのアクター・クリティック強化学習フレームワークを提案する。
  • 外部の社会的ネットワークに依存せず、温暖スタート軌道(WST)からユーザーのネットワーク構造を学習する。
  • グラフラプラシアンを用いて価値関数にℓ₂正則化を施す投影ステップを実装し、コヒーレンスを強制する。
  • パラメータ化されたℓ₂ペナルティを用いて、方策関数に固定点更新を適用し、コヒーレンス制約を課す。
  • Cohesion-RL#1 および Cohesion-RL#2 の2つのアルゴリズムを導入し、価値関数および方策関数の更新におけるℓ₂制約の適用方法に差を設ける。
  • 間接的報酬信号を用いて、価値関数、方策、およびコヒーレンスネットワークを同時に反復的に最適化する。

実験結果

リサーチクエスチョン

  • RQ1類似ユーザー間での情報共有は、データが乏しいmHealth環境下での方策学習を改善できるか?
  • RQ2外部の社会的ネットワークの仮定に依存せず、ユーザー行動データから直接ユーザーネットワークを学習できるか?
  • RQ3異なるコヒーレンス制約の強度が、方策のパフォーマンスと安定性に与える影響は何か?
  • RQ4温暖スタート軌道の長さが、ネットワークコヒーレンスの学習および総合的な方策品質に与える影響は何か?
  • RQ5協調的RLは、オンラインmHealth干渉タスクにおいて、個別方策学習を上回る性能を示せるか?

主な発見

  • Cohesion-RL#2 は、すべてのパrameter設定においてCohesion-RL#1およびSeparate-RLを一貫して上回り、優れた安定性とパフォーマンスを示した。
  • 温暖スタート軌道が短い(T₀ = 5)場合、Cohesion-RL手法により平均ステップ数が261.67から52.09に減少し、ネットワークコヒーレンスの恩恵が顕著に現れた。
  • T₀ を 5 から 20 に増加させることで、Separate-RLのパフォーマンスは210.51ステップ向上し、初期データ品質の重要性が明確に示された。
  • 本手法は限られたデータでも顕著なパフォーマンス向上を達成でき、ユーザーのデータが希薄な現実世界のmHealthシナリオにおいても有効であることが証明された。
  • μ₁ ∈ [0.001, 10] の範囲でネットワークコヒーレンス制約を適用すると、ベースラインよりパフォーマンスが向上し、特に μ₁ = 1.0 の場合が最適な結果を示した。
  • 本手法はパrameterの変動に強く、異なる割引率(γ)に対しても高いパフォーマンスを維持した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。