[Paper Review] Testing the Tools of Systems Neuroscience on Artificial Neural Networks
This paper proposes using artificial neural networks (ANNs) as controlled testbeds to evaluate the effectiveness of systems neuroscience tools in uncovering meaningful insights about neural computation. By applying common analysis tools like demixed PCA to ANNs with known architectures and behaviors, researchers can systematically test whether these tools yield actionable, experimentally confirmable insights—thereby improving the reliability and utility of neuroscience methodologies.
Neuroscientists apply a range of common analysis tools to recorded neural activity in order to glean insights into how neural circuits implement computations. Despite the fact that these tools shape the progress of the field as a whole, we have little empirical evidence that they are effective at quickly identifying the phenomena of interest. Here I argue that these tools should be explicitly tested and that artificial neural networks (ANNs) are an appropriate testing grounds for them. The recent resurgence of the use of ANNs as models of everything from perception to memory to motor control stems from a rough similarity between artificial and biological neural networks and the ability to train these networks to perform complex high-dimensional tasks. These properties, combined with the ability to perfectly observe and manipulate these systems, makes them well-suited for vetting the tools of systems and cognitive neuroscience. I provide here both a roadmap for performing this testing and a list of tools that are suitable to be tested on ANNs. Using ANNs to reflect on the extent to which these tools provide a productive understanding of neural systems -- and on exactly what understanding should mean here -- has the potential to expedite progress in the study of the brain.
Motivation & Objective
- To address the lack of empirical validation for common tools in systems neuroscience, which are widely used but rarely tested for their actual utility in revealing neural mechanisms.
- To argue that artificial neural networks (ANNs) are ideal testbeds due to their biological plausibility, trainability on complex tasks, and full observability and controllability.
- To establish a transparent, pre-registered framework for evaluating analysis tools on ANNs, including both successful and failed outcomes, to avoid publication bias.
- To develop a graded scale for measuring the success of analysis tools, from null/interpretable results to experimentally confirmed actionable insights.
- To encourage the development of better tools by identifying gaps in current methods through systematic testing on ANNs.
Proposed method
- Use ANNs trained on biologically relevant tasks (e.g., image classification with noise) as controlled models of neural circuits with known internal dynamics.
- Apply standard systems neuroscience tools—such as demixed PCA (dPCA), dimensionality reduction, and population vector analysis—to the hidden unit activity of these networks.
- Evaluate tool performance using a four-tiered success metric: null/uninterpretable results, interpretable but non-actionable insights, actionable insights guiding new experiments, and insights confirmed by experimental perturbations.
- Design perturbation experiments based on analysis outputs (e.g., reconstructing activity without components tied to noise) to test whether predicted effects match actual performance changes.
- Pre-register all analysis plans and report all outcomes—positive and negative—using transparent, open documentation to avoid file drawer bias.
- Categorize results based on whether tools successfully demixed task-relevant components (e.g., digit identity vs. noise type) and whether those components predicted functional roles in classification.
Experimental results
Research questions
- RQ1To what extent do standard systems neuroscience tools like dPCA successfully identify interpretable, functionally relevant components in artificial neural networks?
- RQ2Can the insights generated by these tools be experimentally validated through targeted perturbations of network activity?
- RQ3How do different analysis tools compare in their ability to yield actionable, predictive insights about neural network function?
- RQ4What criteria can be used to objectively assess the quality and utility of analysis tools in neural data science?
- RQ5Can systematic testing on ANNs reveal limitations of current tools and guide the development of improved methods?
Key findings
- dPCA applied to an ANN trained on noisy digit classification successfully identified distinct subspaces for digit identity and noise type, indicating effective demixing of task parameters.
- When network activity was reconstructed without dPCs associated with noise, classification performance increased, confirming that these components were not essential for the task, thus validating the analysis insight.
- Analysis tools that yield only interpretable but non-actionable results provide limited value, as they fail to guide testable predictions.
- Tools producing null or uninterpretable results should be documented and reported transparently to avoid misleading conclusions about their utility.
- The highest level of success—where analysis insights lead to predictions that are confirmed by subsequent experiments—was demonstrated in the dPCA perturbation experiment, showing the method’s full potential.
- Pre-registration and full reporting of all test outcomes, including failures, are essential to avoid bias and ensure methodological rigor in tool evaluation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.