Skip to main content
QUICK REVIEW

[论文解读] A Temporal Difference Reinforcement Learning Theory of Emotion: unifying emotion, cognition and adaptive behavior

Joost Broekens|arXiv (Cornell University)|Jul 24, 2018
Cognitive Science and Mapping被引用 4
一句话总结

本文提出了一种基于时间差分强化学习(TDRL)的情绪理论,认为所有情绪均源于大脑对时间差分(TD)误差——即预期奖励与实际奖励之间的差异——的评估,从而将情绪、认知与适应性行为统一起来。该理论整合了心理学、神经生物学和计算证据,解释了情绪如何驱动学习与跨物种的生存导向行为。

ABSTRACT

Emotions are intimately tied to motivation and the adaptation of behavior, and many animal species show evidence of emotions in their behavior. Therefore, emotions must be related to powerful mechanisms that aid survival, and, emotions must be evolutionary continuous phenomena. How and why did emotions evolve in nature, how do events get emotionally appraised, how do emotions relate to cognitive complexity, and, how do they impact behavior and learning? In this article I propose that all emotions are manifestations of reward processing, in particular Temporal Difference (TD) error assessment. Reinforcement Learning (RL) is a powerful computational model for the learning of goal oriented tasks by exploration and feedback. Evidence indicates that RL-like processes exist in many animal species. Key in the processing of feedback in RL is the notion of TD error, the assessment of how much better or worse a situation just became, compared to what was previously expected (or, the estimated gain or loss of utility - or well-being - resulting from new evidence). I propose a TDRL Theory of Emotion and discuss its ramifications for our understanding of emotions in humans, animals and machines, and present psychological, neurobiological and computational evidence in its support.

研究动机与目标

  • 在强化学习基础上,构建一个统一情绪、认知与适应性行为的计算框架。
  • 通过将情绪与动物和人类中基本的奖励处理机制相联系,解释情绪在进化上的连续性。
  • 证明情绪状态源于大脑对奖励预测中时间差分(TD)误差的评估。
  • 提供一种神经生物学上合理的机制,说明情绪如何影响学习、决策与行为。
  • 通过将情绪建模为目标导向学习的内在组成部分,弥合人工智能与情感科学之间的鸿沟。

提出的方法

  • 提出情绪状态对应于强化学习中时间差分(TD)误差的大小与效价。
  • 将情绪评估建模为大脑对奖励预测误差的实时评估——即结果比预期好或差多少。
  • 使用时间差分学习方程,特别是TD(0)更新规则:δ = r + γV(s′) − V(s),以形式化情绪的效价与强度。
  • 整合神经生物学证据,表明多巴胺等神经调质编码与情绪状态一致的时间差分误差信号。
  • 将该框架应用于解释多种情绪现象,如惊讶、愉悦、恐惧与挫败感,这些均可视为时间差分误差处理的不同表现。
  • 将模型扩展,以说明情绪反馈如何通过基于奖励的学习,塑造长期行为与认知发展。

实验结果

研究问题

  • RQ1情绪如何从大脑的基本奖励处理机制中产生?
  • RQ2情绪在引导适应性行为与学习中的计算作用是什么?
  • RQ3如何利用强化学习原理正式建模情绪状态?
  • RQ4强化学习中的时间差分误差在何种方式上对应于主观情绪体验?
  • RQ5该理论如何解释情绪在不同物种间的进化连续性?

主要发现

  • 情绪并非与认知分离,而是奖励学习的内在组成部分,情绪效价反映时间差分误差的符号。
  • 情绪反应的强度与时间差分误差的大小相关,解释了对意外奖励或惩罚产生更强反应的原因。
  • 神经生物学证据表明,多巴胺等神经调质编码时间差分误差,与愉悦或厌恶等情绪状态一致。
  • 该理论可直接解释多种情绪现象,包括惊讶、挫败感与狂喜,均为时间差分误差计算的直接结果。
  • 情绪反馈通过调节注意力、探索与决策,提升学习效率,与观察到的行为适应一致。
  • 该模型提供了一个统一框架,通过单一计算机制——时间差分误差处理,解释人类与动物的情绪行为。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。