[Paper Review] Training recurrent neural networks with sparse, delayed rewards for flexible decision tasks
This paper proposes a biologically plausible, reward-modulated Hebbian learning rule to train recurrent neural networks (RNNs) using only sparse, delayed rewards—without real-time error signals or dedicated feedback networks. The method successfully trains RNNs on a delayed nonmatch-to-sample task, producing dynamic, task-relevant neural representations that evolve from stimulus- to response-specific coding, mirroring cortical activity in behaving animals.
Recurrent neural networks in the chaotic regime exhibit complex dynamics reminiscent of high-level cortical activity during behavioral tasks. However, existing training methods for such networks are either biologically implausible, or require a real-time continuous error signal to guide the learning process. This is in contrast with most behavioral tasks, which only provide time-sparse, delayed rewards. Here we show that a biologically plausible reward-modulated Hebbian learning algorithm, previously used in feedforward models of birdsong learning, can train recurrent networks based solely on delayed, phasic reward signals at the end of each trial. The method requires no dedicated feedback or readout networks: the whole network connectivity is subject to learning, and the network output is read from one arbitrarily chosen network cell. We use this method to successfully train a network on a delayed nonmatch to sample task (which requires memory, flexible associations, and non-linear mixed selectivities). Using decoding techniques, we show that the resulting networks exhibit dynamic coding of task-relevant information, with neural encodings of various task features fluctuating widely over the course of a trial. Furthermore, network activity moves from a stimulus-specific representation to a response-specific representation during response time, in accordance with neural recordings in behaving animals for similar tasks. We conclude that recurrent neural networks, trained with reward-modulated Hebbian learning, offer a plausible model of cortical dynamics during learning and performance of flexible association.
Motivation & Objective
- To develop a biologically plausible learning rule for training recurrent neural networks without continuous error signals.
- To enable RNNs to learn complex, flexible decision tasks using only sparse, phasic reward signals at the end of each trial.
- To train networks without dedicated feedback or readout networks, allowing all connections to be plastic.
- To investigate whether such networks can exhibit dynamic, task-relevant neural coding similar to that observed in cortical recordings.
- To demonstrate that reward-modulated Hebbian learning can produce networks capable of memory, non-linear associations, and flexible behavior.
Proposed method
- The method employs a reward-modulated Hebbian learning rule, where synaptic weight changes depend on the product of presynaptic activity, postsynaptic activity, and a delayed reward signal.
- The network is trained on a delayed nonmatch-to-sample task requiring memory and flexible associations, with rewards given only at the end of each trial.
- The output is read from a single, arbitrarily chosen neuron in the network, eliminating the need for a separate readout or feedback mechanism.
- All network connections, including recurrent ones, are subject to learning, enabling full end-to-end training with only reward feedback.
- Decoding techniques are used to analyze the dynamic representation of task features across time, revealing changes in neural coding.
- The learning rule is implemented without backpropagation or real-time error signals, making it compatible with biological plausibility.
Experimental results
Research questions
- RQ1Can a biologically plausible, reward-modulated Hebbian learning rule train recurrent neural networks to perform complex, flexible decision tasks with only delayed, sparse rewards?
- RQ2Does the resulting network exhibit dynamic neural coding that evolves from stimulus-specific to response-specific representations during a trial?
- RQ3Can the network learn non-linear mixed selectivities and maintain memory over delays without dedicated feedback or readout networks?
- RQ4How do the network's internal representations compare to those observed in cortical recordings during similar behavioral tasks?
- RQ5Is it possible to achieve flexible association learning in RNNs using only phasic reward signals and no continuous error signals?
Key findings
- The network successfully learned the delayed nonmatch-to-sample task using only sparse, phasic reward signals at the end of each trial.
- Neural activity dynamically encoded task-relevant information, shifting from stimulus-specific to response-specific representations over the course of a trial.
- The network exhibited non-linear mixed selectivities and maintained memory across delays, demonstrating flexible decision-making.
- Decoding analysis confirmed that task features were represented in fluctuating, time-varying patterns across the network's activity.
- The learning rule enabled end-to-end training without feedback or readout networks, with all connections plastic and trained via reward-modulated Hebbian plasticity.
- The resulting network dynamics closely resembled those observed in cortical recordings from behaving animals performing similar tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.