[Paper Review] Deep Learning Models of the Retinal Response to Natural Scenes
Convolutional neural networks accurately predict retinal ganglion cell responses to natural scenes, outperform LN/GLMs, generalize across stimulus types, and reveal internal retinal mechanisms.
A central challenge in neuroscience is to understand neural computations and circuit mechanisms that underlie the encoding of ethologically relevant, natural stimuli. In multilayered neural circuits, nonlinear processes such as synaptic transmission and spiking dynamics present a significant obstacle to the creation of accurate computational models of responses to natural stimuli. Here we demonstrate that deep convolutional neural networks (CNNs) capture retinal responses to natural scenes nearly to within the variability of a cell's response, and are markedly more accurate than linear-nonlinear (LN) models and Generalized Linear Models (GLMs). Moreover, we find two additional surprising properties of CNNs: they are less susceptible to overfitting than their LN counterparts when trained on small amounts of data, and generalize better when tested on stimuli drawn from a different distribution (e.g. between natural scenes and white noise). Examination of trained CNNs reveals several properties. First, a richer set of feature maps is necessary for predicting the responses to natural scenes compared to white noise. Second, temporally precise responses to slowly varying inputs originate from feedforward inhibition, similar to known retinal mechanisms. Third, the injection of latent noise sources in intermediate layers enables our model to capture the sub-Poisson spiking variability observed in retinal ganglion cells. Fourth, augmenting our CNNs with recurrent lateral connections enables them to capture contrast adaptation as an emergent property of accurately describing retinal responses to natural scenes. These methods can be readily generalized to other sensory modalities and stimulus ensembles. Overall, this work demonstrates that CNNs not only accurately capture sensory circuit responses to natural scenes, but also yield information about the circuit's internal structure and function.
Motivation & Objective
- Understand how retinal ganglion cells encode natural scene stimuli.
- Evaluate CNNs as predictive models for retinal responses to natural scenes versus LN and GLM baselines.
- Investigate generalization across stimulus distributions (natural scenes vs white noise).
- Identify internal retinal-like mechanisms captured by CNNs (inhibition, adaptation, variability).
- Explore architectural augmentations (recurrent connections) to model long timescale dynamics.
Proposed method
- Train deep CNNs to predict ganglion cell spiking from natural scene sequences and white-noise stimuli.
- Compare CNNs against linear-nonlinear (LN) and generalized linear models (GLMs).
- Optimize with Poisson negative log-likelihood loss using ADAM; apply L2 and L1 regularization.
- Vary network depth, filter sizes (>15x15), and layer types; evaluate on held-out data.
- Visualize learned first- and second-layer receptive fields to interpret features.
- Optionally augment CNNs with recurrent layers to capture longer timescale adaptation.
Experimental results
Research questions
- RQ1Can CNNs outperform LN/GLM models in predicting retinal responses to natural scenes?
- RQ2Do CNNs generalize better across stimulus distributions (natural scenes vs white noise)?
- RQ3What internal mechanisms of the retina emerge from CNN representations (e.g., feedforward inhibition, adaptation, sub-Poisson variability)?
- RQ4Do recurrent connections improve modeling of long-timescale adaptive dynamics?
- RQ5How do learned features differ between natural-scene and white-noise training data?
Key findings
- CNNs substantially outperform LN models and GLMs in predicting retinal responses to both natural scenes and white noise.
- CNNs achieve reliability approaching the retina and show better generalization across stimulus distributions than simpler models.
- Training with injected latent noise captures sub-Poisson variability observed in retinal spiking.
- CNNs reveal temporally precise firing via feedforward inhibition and demonstrate broader, more diverse second-layer features for natural scenes.
- Augmenting CNNs with recurrent lateral connections enables emergent contrast adaptation as a property of accurate response description.
- CNNs trained on one stimulus class generalize better to other stimulus classes than GLMs or LN models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.