Skip to main content
QUICK REVIEW

[論文レビュー] A Temporal Difference Reinforcement Learning Theory of Emotion: unifying emotion, cognition and adaptive behavior

Joost Broekens|arXiv (Cornell University)|Jul 24, 2018
Cognitive Science and Mapping被引用数 4
ひとこと要約

本論文は、報酬予測誤差(期待と実際の報酬の差)としての時系列差分(TD)誤差の脳による評価に基づき、すべての感情が生じることを主張する、感情のための時系列差分強化学習(TDRL)理論を提唱する。この理論により、感情、認知、適応的行動が統合される。心理学的・神経生物学的・計算的証拠を統合し、感情が学習および種を越えた生存指向行動を駆動する仕組みを説明する。

ABSTRACT

Emotions are intimately tied to motivation and the adaptation of behavior, and many animal species show evidence of emotions in their behavior. Therefore, emotions must be related to powerful mechanisms that aid survival, and, emotions must be evolutionary continuous phenomena. How and why did emotions evolve in nature, how do events get emotionally appraised, how do emotions relate to cognitive complexity, and, how do they impact behavior and learning? In this article I propose that all emotions are manifestations of reward processing, in particular Temporal Difference (TD) error assessment. Reinforcement Learning (RL) is a powerful computational model for the learning of goal oriented tasks by exploration and feedback. Evidence indicates that RL-like processes exist in many animal species. Key in the processing of feedback in RL is the notion of TD error, the assessment of how much better or worse a situation just became, compared to what was previously expected (or, the estimated gain or loss of utility - or well-being - resulting from new evidence). I propose a TDRL Theory of Emotion and discuss its ramifications for our understanding of emotions in humans, animals and machines, and present psychological, neurobiological and computational evidence in its support.

研究の動機と目的

  • 強化学習に根ざした単一の計算的枠組みとして、感情、認知、適応的行動を統合すること。
  • 動物および人間における基本的な報酬処理メカニズムと関連づけることで、感情の進化的連続性を説明すること。
  • 感情状態が報酬予測における時系列差分(TD)誤差の評価によって生じることを示すこと。
  • 感情が学習、意思決定、行動に与える影響を神経生物学的に妥当なメカニズムとして提供すること。
  • 人工知能と感情科学を結ぶために、感情を目的指向的学習の内在的要素としてモデル化すること。

提案手法

  • 感情状態が強化学習における時系列差分(TD)誤差の大きさと価値(valence)に対応すると提唱する。
  • 感情的評価を、脳による報酬予測誤差のリアルタイム評価としてモデル化する—予想よりどれほど良いか、または悪いかの程度。
  • 特にTD(0)更新則:δ = r + γV(s′) − V(s) を用いて、感情の価値と強度を形式化する。
  • ドーパミンや他の神経調節物質がTD誤差信号を符号化している神経生物学的証拠を統合し、快楽や嫌悪といった感情状態と一致させる。
  • 驚き、快楽、恐怖、いらだちといった多様な感情的現象を、TD誤差処理の変種として説明するためにフレームワークを適用する。
  • 感情フィードバックが報酬ベースの学習を通じて長期的な行動と認知発達に与える影響を示すために、モデルを拡張する。

実験結果

リサーチクエスチョン

  • RQ1脳における基本的な報酬処理メカニズムからどのように感情が生じるのか?
  • RQ2感情が適応的行動と学習を誘導する上で果たす計算的役割は何か?
  • RQ3強化学習の原則を用いて感情状態を形式的にモデル化できるか?
  • RQ4強化学習におけるTD誤差が主観的な感情的体験とどのように対応するか?
  • RQ5この理論は、種を越えて共通する感情の進化的連続性をどのように説明するか?

主な発見

  • 感情は認知とは別個のものではなく、報酬ベースの学習の本質的側面であり、感情の価値はTD誤差の符号に対応する。
  • 感情反応の強度はTD誤差の大きさと相関しており、予期しない報酬や罰に対する強い反応を説明する。
  • ドーパミンや他の神経調節物質がTD誤差を符号化している神経生物学的証拠が、快楽や嫌悪といった感情状態と一致する。
  • 驚き、いらだち、喜びといった多様な感情的現象が、TD誤差計算の直接的結果として説明可能である。
  • 感情フィードバックが注意、探査行動、意思決定を調節することで、学習効率が向上し、観察された行動的適応と一致する。
  • このモデルは、単一の計算的メカニズム—TD誤差処理—を通じて、人間および動物の両方の感情的行動を統一的に説明するフレームワークを提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。