Skip to main content
QUICK REVIEW

[Paper Review] Self-supervised learning through the eyes of a child

A. Emin Orhan, Vaibhav V. Gupta|arXiv (Cornell University)|Jul 31, 2020
Innovative Teaching and Learning MethodsPsychology42 references65 citations
TL;DR

The paper demonstrates that powerful, high-level visual representations can emerge from developmentally realistic egocentric video via self-supervised learning, using a novel temporal classification objective trained on SAYCam data from individual children.

ABSTRACT

Within months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives.

Motivation & Objective

  • Motivate understanding of how much early visual knowledge can arise from generic learning on sensory data.
  • Leverage developmentally informed, longitudinal egocentric video to study representation learning without explicit labels.
  • Assess whether self-supervised learning yields transferable, high-level visual categories relevant to a child’s environment.

Proposed method

  • Train self-supervised deep convolutional networks (MobileNetV2) from scratch on raw, unlabeled headcam videos from individual children.
  • Introduce a temporal classification objective that predicts which episode (temporal class) a frame belongs to, enforcing invariance to fast-changing low-level details.
  • Compare temporal classification with static and temporal contrastive learning baselines on downstream tasks.
  • Evaluate learned representations by freezing the trunk and training linear readouts on developmentally relevant categories.
  • Use a curated labeled subset of SAYCam data (from one child) and the Toybox dataset to assess generalization and robustness.

Experimental results

Research questions

  • RQ1Can generic self-supervised learning on developmentally realistic, longitudinal egocentric video yield high-level visual representations?
  • RQ2Do temporal-invariance based learning objectives outperform image-based or contrastive objectives on downstream child-relevant categorization tasks?
  • RQ3To what extent do learned representations generalize across children and to unseen exemplars?
  • RQ4What factors (sampling rate, segment length, data augmentation) influence downstream task performance?
  • RQ5Are the learned features localized and behaviorally plausible for visual categorization in a child’s environment?

Key findings

  • Temporal classification self-supervised models achieve high downstream accuracy on labeled child data and Toybox tasks, sometimes rivaling ImageNet-pretrained baselines.
  • Temporal models trained on data from different children generalize to labeled data from another child.
  • Temporal classification outperforms static contrastive and temporal contrastive learners in all reported conditions.
  • Learned representations show invariance to natural transformations and can generalize to unseen exemplars with limited labeled data.
  • Analyses indicate distributed feature representations with higher selectivity in higher layers, and attention maps align with meaningful image regions for certain categories.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.