[论文解读] Action-modulated midbrain dopamine activity arises from distributed control policies
本文提出了一种基底神经节中生物上合理的离策略强化学习模型,通过引入‘动作意外’项(衡量动作与基底神经节预期动作的偏差)来解释中脑多巴胺活动,该模型结合了奖励预测误差(RPE)。该模型能够有效学习由其他脑区(如运动皮层、小脑)驱动的行为,而经典在策略模型则无法实现,且可解释实验发现,如与运动相关的多巴胺信号、随练习减少的调制效应,以及背侧/腹侧纹状体活动的差异。
Animal behavior is driven by multiple brain regions working in parallel with distinct control policies. We present a biologically plausible model of off-policy reinforcement learning in the basal ganglia, which enables learning in such an architecture. The model accounts for action-related modulation of dopamine activity that is not captured by previous models that implement on-policy algorithms. In particular, the model predicts that dopamine activity signals a combination of reward prediction error (as in classic models) and "action surprise," a measure of how unexpected an action is relative to the basal ganglia's current policy. In the presence of the action surprise term, the model implements an approximate form of Q-learning. On benchmark navigation and reaching tasks, we show empirically that this model is capable of learning from data driven completely or in part by other policies (e.g. from other brain regions). By contrast, models without the action surprise term suffer in the presence of additional policies, and are incapable of learning at all from behavior that is completely externally driven. The model provides a computational account for numerous experimental findings about dopamine activity that cannot be explained by classic models of reinforcement learning in the basal ganglia. These include differing levels of action surprise signals in dorsal and ventral striatum, decreasing amounts movement-modulated dopamine activity with practice, and representations of action initiation and kinematics in dopamine activity. It also provides further predictions that can be tested with recordings of striatal dopamine activity.
研究动机与目标
- 解决经典在策略强化学习模型在基底神经节中的局限性,这些模型在行为由外部控制器(如运动皮层、小脑)驱动时会失效。
- 解释为何中脑多巴胺神经元表现出超出奖励预测误差(RPE)的与动作相关的信号,如动作启动与强度。
- 在单一生物上合理的学习框架内,统一整合RPE与动作意外,提供对多巴胺活动的计算解释。
- 证明动作意外对于从离策略行为中学习至关重要,尤其是在复杂、多阶段任务中。
提出的方法
- 引入一种修改后的多巴胺信号,结合奖励预测误差(RPE)与动作意外项:||a − μ(s)||²,其中a为实际动作,μ(s)为基底神经节在状态s下的预期动作。
- 将算法推导为连续动作空间中Q-learning的近似形式,实现离策略学习。
- 采用策略梯度框架,其中皮层-纹状体突触的可塑性由RPE与动作意外信号的联合信号调节。
- 采用一种生物上合理的学习规则,其中多巴胺释放根据结果预测误差与动作偏离预期行为的程度,调控可塑性。
- 通过外部策略(如运动皮层)生成的数据,模拟导航与抓握任务中的学习过程,结果表明其性能优于在策略模型。
- 对背侧与腹侧纹状体中动作意外表达的差异进行建模,表明这种不对称性可加速外部控制器策略的学习。
实验结果
研究问题
- RQ1当行为由多个具有不同控制策略的脑区驱动(而非基底神经节自身策略)时,基底神经节如何实现有效学习?
- RQ2为何中脑多巴胺神经元表现出RPE无法解释的与动作相关的活动?
- RQ3动作意外在多巴胺信号传导中的功能角色是什么?它如何支持基底神经节中的离策略学习?
- RQ4该模型如何解释随练习而减少的运动调制多巴胺活动?
- RQ5该模型能否解释动作意外在背侧与腹侧纹状体中的差异性表征?
主要发现
- 引入动作意外的模型能够有效学习完全由外部控制器(如运动皮层)驱动的行为,而经典在策略模型在此类条件下完全失效。
- 在需要长序列动作的任务中,动作意外对学习至关重要,因为基底神经节与外部控制器之间的策略不匹配会导致标准RPE模型失败。
- 该模型可解释实验中观察到的与运动相关的多巴胺活动,包括与动作启动和运动学相关的信号,而无需引入独立的动机功能。
- 该模型预测背侧纹状体多巴胺神经元的行动意外信号强于腹侧纹状体神经元,与实验数据中背侧区域运动调制更强的现象一致。
- 该模型解释了随练习而减少的运动调制多巴胺活动,其原因是行为变得更具可预测性,导致动作意外降低。
- 该模型提供了一种机制,使动作意外可作为监督学习信号,加速外部控制器策略在基底神经节中的整合与巩固。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。