[论文解读] Training recurrent neural networks with sparse, delayed rewards for flexible decision tasks
该论文提出了一种生物上合理的、受奖励调制的赫布可塑性学习规则,用于仅通过稀疏、延迟的奖励来训练循环神经网络(RNNs),而无需实时误差信号或专用反馈网络。该方法成功地在延迟非匹配到样本任务上训练了RNN,产生了从刺激相关到响应相关的动态、任务相关的神经表征,其演变过程与行为动物皮层活动相似。
Recurrent neural networks in the chaotic regime exhibit complex dynamics reminiscent of high-level cortical activity during behavioral tasks. However, existing training methods for such networks are either biologically implausible, or require a real-time continuous error signal to guide the learning process. This is in contrast with most behavioral tasks, which only provide time-sparse, delayed rewards. Here we show that a biologically plausible reward-modulated Hebbian learning algorithm, previously used in feedforward models of birdsong learning, can train recurrent networks based solely on delayed, phasic reward signals at the end of each trial. The method requires no dedicated feedback or readout networks: the whole network connectivity is subject to learning, and the network output is read from one arbitrarily chosen network cell. We use this method to successfully train a network on a delayed nonmatch to sample task (which requires memory, flexible associations, and non-linear mixed selectivities). Using decoding techniques, we show that the resulting networks exhibit dynamic coding of task-relevant information, with neural encodings of various task features fluctuating widely over the course of a trial. Furthermore, network activity moves from a stimulus-specific representation to a response-specific representation during response time, in accordance with neural recordings in behaving animals for similar tasks. We conclude that recurrent neural networks, trained with reward-modulated Hebbian learning, offer a plausible model of cortical dynamics during learning and performance of flexible association.
研究动机与目标
- 开发一种无需连续误差信号的、生物上合理的RNN训练规则。
- 使RNN能够仅通过每轮试验结束时的稀疏、突触状奖励信号,学习复杂且灵活的决策任务。
- 在无需专用反馈或读出网络的情况下训练网络,使所有连接均可塑。
- 研究此类网络是否能表现出类似于在皮层记录中观察到的动态、任务相关的神经编码。
- 证明奖励调制的赫布可塑性学习能够产生具备记忆、非线性关联和灵活行为能力的网络。
提出的方法
- 该方法采用一种受奖励调制的赫布可塑性学习规则,其中突触权重的变化取决于突触前活动、突触后活动与延迟奖励信号的乘积。
- 网络在延迟非匹配到样本任务上进行训练,该任务要求具备记忆功能和灵活关联能力,奖励仅在每轮试验结束时给予。
- 输出从网络中一个任意选择的神经元读取,从而无需单独的读出或反馈机制。
- 所有网络连接(包括循环连接)均参与学习,实现了仅通过奖励反馈的端到端训练。
- 采用解码技术分析任务特征在时间上的动态表征,揭示了神经编码的变化。
- 学习规则的实现不依赖反向传播或实时误差信号,因此具备生物合理性。
实验结果
研究问题
- RQ1一种生物上合理的、受奖励调制的赫布可塑性学习规则,能否仅通过稀疏、延迟的奖励信号,训练循环神经网络完成复杂且灵活的决策任务?
- RQ2由此产生的网络是否表现出动态神经编码,其在试验过程中从刺激特异性表征逐渐演变为响应特异性表征?
- RQ3该网络能否在无需专用反馈或读出网络的情况下,学习非线性混合选择性并维持延迟期的记忆?
- RQ4该网络的内部表征与在类似行为任务中观察到的皮层记录中的表征相比如何?
- RQ5是否可能仅通过突触状奖励信号且无连续误差信号,实现在RNN中的灵活关联学习?
主要发现
- 网络仅通过每轮试验结束时的稀疏、突触状奖励信号,成功学习了延迟非匹配到样本任务。
- 神经活动动态地编码了任务相关的信息,其表征在试验过程中从刺激特异性逐渐转变为响应特异性。
- 该网络表现出非线性混合选择性,并在延迟期保持了记忆,展示了灵活的决策能力。
- 解码分析证实,任务特征以随时间波动的、动态变化的模式在全网络活动中得到表征。
- 学习规则实现了无需反馈或读出网络的端到端训练,所有连接均可塑,并通过奖励调制的赫布可塑性进行训练。
- 由此产生的网络动力学与行为动物在执行类似任务时的皮层记录结果高度相似。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。