[Paper Review] Controlled-rearing studies of newborn chicks and deep neural networks
This study directly compares visual learning in newborn chicks and deep convolutional neural networks (CNNs) using parallel controlled-rearing experiments. By simulating the same visual environments for both chicks and CNNs—using a video game engine to generate training data from a chick’s perspective—it demonstrates that CNNs achieve view-invariant object recognition with the same sparse data that enables rapid learning in chicks, challenging the notion that CNNs are inherently 'data hungry.'
Convolutional neural networks (CNNs) can now achieve human-level performance on challenging object recognition tasks. CNNs are also the leading quantitative models in terms of predicting neural and behavioral responses in visual recognition tasks. However, there is a widely accepted critique of CNN models: unlike newborn animals, which learn rapidly and efficiently, CNNs are thought to be "data hungry," requiring massive amounts of training data to develop accurate models for object recognition. This critique challenges the promise of using CNNs as models of visual development. Here, we directly examined whether CNNs are more data hungry than newborn animals by performing parallel controlled-rearing experiments on newborn chicks and CNNs. We raised newborn chicks in strictly controlled visual environments, then simulated the training data available in that environment by constructing a virtual animal chamber in a video game engine. We recorded the visual images acquired by an agent moving through the virtual chamber and used those images to train CNNs. When CNNs received similar visual training data as chicks, the CNNs successfully solved the same challenging view-invariant object recognition tasks as the chicks. Thus, the CNNs were not more data hungry than animals: both CNNs and chicks successfully developed robust object models from training data of a single object.
Motivation & Objective
- To test whether deep convolutional neural networks (CNNs) are truly 'data hungry' compared to newborn animals, particularly in view-invariant object recognition.
- To investigate whether CNNs can achieve similar learning efficiency as newborn chicks when trained on the same limited visual input.
- To develop a computational model of chick visual experience using a video game engine to simulate real-world training data for CNNs.
- To enable a direct, fair comparison between biological and artificial visual systems by aligning training data and testing protocols.
- To evaluate whether CNNs can serve as neurally mechanistic, image-computable models of early visual development in animals.
Proposed method
- Conducted controlled-rearing experiments with newborn chicks in a virtual environment with a single object, measuring their performance on a view-invariant object recognition task.
- Built a virtual animal chamber in a video game engine to simulate the chick's visual environment and record first-person visual trajectories from a moving agent.
- Collected and processed raw visual images from the agent’s movement through the virtual chamber to create a training dataset for CNNs.
- Trained both supervised and unsupervised CNNs on the simulated training data, using data augmentation to increase effective sample size.
- Used linear discriminant analysis (LDA) to visualize feature representations and assess representational similarity across models.
- Tested both chicks and CNNs on the same object recognition tasks to enable direct comparison of learning performance and generalization.
Experimental results
Research questions
- RQ1Can deep convolutional neural networks (CNNs) achieve view-invariant object recognition when trained on the same minimal visual data that enables rapid learning in newborn chicks?
- RQ2How does the amount and quality of visual training data from a single-object environment compare between newborn chicks and CNNs in terms of learning efficiency?
- RQ3To what extent do CNNs replicate the representational and behavioral outcomes observed in newborn chicks during early visual development?
- RQ4What role does data augmentation and active exploration play in reducing data requirements for both biological and artificial visual systems?
- RQ5Can CNNs serve as valid, neurally mechanistic models of early visual development in animals when trained on real-world-like sensory input?
Key findings
- CNNs trained on visual data collected from a virtual agent moving through the same controlled-rearing environment as newborn chicks successfully solved the same view-invariant object recognition task.
- Both chicks and CNNs developed robust, view-invariant object representations after exposure to a single object, indicating comparable learning efficiency.
- The number of effective training images for CNNs was significantly increased through data augmentation, making the data load comparable to that of newborn chicks, who acquire ~864,000 training samples in 24 hours.
- Representational dissimilarity matrices (RDMs) showed that CNNs developed feature representations similar in structure to those observed in biological systems.
- The study challenges the widely held belief that CNNs are inherently 'data hungry,' showing that with aligned training data, they perform as efficiently as newborn animals.
- The results support the use of CNNs as image-computable, neurally mechanistic models of early visual development, offering a framework for testing and falsifying developmental theories.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.