Skip to main content
QUICK REVIEW

[Paper Review] Towards Robust Evaluations of Continual Learning

Sebastian Farquhar, Yarin Gal|arXiv (Cornell University)|May 24, 2018
Domain Adaptation and Few-Shot LearningComputer Science38 references185 citations
TL;DR

This paper critiques current continual learning evaluations, proposes core desiderata for robust evaluation, analyzes biases toward prior-focused methods, and introduces improved experimental designs to better reflect real-world continual learning challenges.

ABSTRACT

Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups. We examine standard evaluations and show why these evaluations make some continual learning approaches look better than they are. We introduce desiderata for continual learning evaluations and explain why their absence creates misleading comparisons. Based on our desiderata we then propose new experiment designs which we demonstrate with various continual learning approaches and datasets. Our analysis calls for a reprioritization of research effort by the community.

Motivation & Objective

  • Define a formal and motivational framework for continual learning evaluations.
  • Identify and critique common evaluation flaws that bias toward prior-focused methods.
  • Propose a set of core desiderata for robust continual learning benchmarks applicable across datasets.
  • Demonstrate that prior-focused methods falter under comprehensive, realistic evaluation regimes.
  • Introduce new experimental designs that address identified shortcomings and better reflect continual learning challenges.

Proposed method

  • Formalize continual learning as sequential task learning with non-i.i.d. data splits.
  • Categorize approaches into prior-focused, likelihood-focused, and hybrid methods with Bayesian interpretations.
  • Critically analyze common evaluation setups (Permuted MNIST, Split MNIST, two-task transfer) for their alignment with real-world needs.
  • Propose five core desiderata for evaluations: cross-task resemblance, shared output head, no test-time task labels, no unconstrained retraining, and scalability to many tasks.
  • Empirically compare representative methods (VCL, EWC, VGR) under evaluations that satisfy all desiderata versus subset evaluations.
  • Recommend and illustrate new evaluation designs that better capture time/memory constraints and privacy considerations.

Experimental results

Research questions

  • RQ1Do common continual learning evaluations faithfully reflect core continual learning challenges?
  • RQ2Are priors-focused methods biased by standard evaluation setups, and under what designs do they fail?
  • RQ3What core desiderata should guide robust continual learning evaluations across datasets?
  • RQ4Can new experimental designs mitigate biases and reveal fundamental limitations of current approaches?
  • RQ5How do time, memory, and privacy considerations integrate into robust continual learning benchmarks?

Key findings

  • Many leading prior-focused methods perform poorly under evaluations that satisfy all core desiderata, revealing blind spots not seen in standard benchmarks.
  • Evaluation setups like Permuted MNIST and multi-headed Split MNIST can bias results in favor of prior-focused approaches.
  • Likelihood-focused methods (e.g., VGR) tend to be more robust when evaluated under comprehensive desiderata-based designs.
  • Two-task transfers and simplistic datasets do not capture long-horizon continual learning challenges, potentially overestimating method capabilities.
  • Proposed evaluation designs expose trade-offs in time/accuracy and incorporate model uncertainty as a tool for task boundary detection.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.