[Paper Review] Modulating human brain responses via optimal natural image selection and synthetic image generation
This study introduces NeuroGen, a data-driven framework that uses deep generative models and personalized neural encoding models to design synthetic and natural images that optimally modulate regional brain activity in fMRI. It demonstrates that synthetic images tailored via group-level and individual-level models elicit significantly stronger responses in specific visual regions—especially aTLfaces and FBA1—compared to natural images, with personalized models outperforming group models in individual-specific activation gains.
One of the main goals of neuroscience is to understand how biological brains interpret and process incoming environmental information. Building computational encoding models that map images to neural responses is one way to pursue this goal. Moreover, generating or selecting visual stimuli designed to achieve specific patterns of responses allows exploration and control of neuronal firing rates or regional brain activity responses. Here, we investigated the brain's regional activation selectivity and inter-individual differences in human brain responses to various sets of natural and synthetic (generated) images via two functional MRI (fMRI) studies. For our first fMRI study, we used a pre-trained group-level neural model for selecting or synthesizing images that are predicted to maximally activate targeted brain regions. We then presented these images to subjects while collecting their fMRI data. Our results show that optimized images indeed evoke larger magnitude responses than other images predicted to achieve average levels of activation.Furthermore, the activation gain is positively associated with the encoding model accuracy. While most regions' activations in response to maximal natural images and maximal synthetic images were not different, two regions, namely anterior temporal lobe faces (aTLfaces) and fusiform body area 1 (FBA1), had significantly higher activation in response to maximal synthetic images compared to maximal natural images. On the other hand, three regions; medial temporal lobe face area (mTLfaces), ventral word form area 1 (VWFA1) and ventral word form area 2 (VWFA2), had higher activation in response to maximal natural images compared to maximal synthetic images. In our second fMRI experiment, we focused on probing inter-individual differences in face regions' responses and found that individual-specific synthetic (and not natural) images derived using a personalized encoding model elicited significantly higher responses compared to synthetic images derived from the group-level or other subjects' encoding models. Finally, we replicated the finding showing synthetic images elicited larger activation responses in the aTLfaces region compared to natural image responses in that region. Here, for the first time, we leverage our data-driven and generative modeling framework NeuroGen to probe inter-individual differences in and functional specialization of the human visual system. Our results indicate that NeuroGen can be used to modulate macro-scale brain regions in specific individuals using synthetically generated visual stimuli.
Motivation & Objective
- To develop a framework for generating visual stimuli that optimally activate targeted brain regions using deep generative models and neural encoding models.
- To investigate whether synthetic images can elicit stronger fMRI responses than natural images in specific human visual cortical regions.
- To explore inter-individual differences in brain response patterns and assess whether personalized encoding models improve stimulus design for individual subjects.
- To validate that synthetic stimuli designed via optimal image generation can reliably modulate macro-scale brain activity in a controlled, data-driven manner.
Proposed method
- Trained subject-specific and group-level deep neural network (DNN)-based encoding models using fMRI data from the Natural Scenes Dataset (NSD), with ridge regression mapping image features to regional brain responses.
- Constructed personalized encoding models via linear ensemble learning, combining base models trained on individual NSD subjects’ data using small, prospective data from Session 1.
- Employed the NeuroGen framework, which couples a pre-trained BigGAN-deep generator with an encoding model, optimizing the noise vector to minimize a loss function that matches desired brain activation patterns.
- For 'Max' conditions, the loss was the negative predicted activation plus L2 regularization on the noise vector; for 'Avg' conditions, it was the absolute difference from average activation.
- Used linear mixed-effects (LME) models with permutation testing to assess statistical significance of response differences across image conditions, accounting for subject-specific random effects.
- Applied Gaussian pooling to reduce feature dimensionality before final ridge regression, improving model efficiency and generalization.
Experimental results
Research questions
- RQ1Can synthetic images generated via a deep generative model and encoding model framework elicit stronger fMRI responses than natural images in targeted brain regions?
- RQ2Do individual-specific synthetic stimuli, derived from personalized encoding models, produce greater activation than stimuli derived from group-level models in the same individuals?
- RQ3Are there regional differences in the brain’s response to synthetic versus natural images, and which regions show preferential sensitivity to synthetic stimuli?
- RQ4How does the accuracy of the underlying encoding model correlate with the magnitude of activation gain in response to optimized stimuli?
Key findings
- Synthetic images optimized via the group-level encoding model elicited significantly higher fMRI responses than average-predicted images in all tested brain regions, confirming the efficacy of the optimization framework.
- In the anterior temporal lobe face area (aTLfaces) and fusiform body area 1 (FBA1), synthetic images produced significantly higher activation than the most activating natural images, indicating synthetic stimuli can surpass natural stimuli in evoking responses.
- Conversely, in the medial temporal lobe face area (mTLfaces), ventral word form area 1 (VWFA1), and VWFA2, natural images elicited significantly higher responses than synthetic images, highlighting region-specific differences in stimulus preference.
- Personalized synthetic stimuli, derived from individual-specific encoding models, elicited significantly higher fMRI responses than those derived from group-level or other subjects’ models, demonstrating the value of personalization.
- The activation gain from optimized stimuli was positively correlated with the accuracy of the underlying encoding model, suggesting model fidelity predicts stimulus effectiveness.
- The study replicated prior findings that synthetic stimuli can evoke larger responses than natural stimuli in the aTLfaces region, now with a controlled, generative framework and individual-level validation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.