Skip to main content
QUICK REVIEW

[Paper Review] Compute Trends Across Three Eras of Machine Learning

Jaime Sevilla, Lennart Heim|arXiv (Cornell University)|Feb 11, 2022
Machine Learning and Data ClassificationComputer Science112 references40 citations
TL;DR

The paper analyzes how training compute (FLOPs) has evolved across three eras—Pre Deep Learning, Deep Learning, and Large-Scale—finding distinct doubling times and a late emergence of large-scale models with a separate trend.

ABSTRACT

Compute, data, and algorithmic advances are the three fundamental factors that guide the progress of modern Machine Learning (ML). In this paper we study trends in the most readily quantified factor - compute. We show that before 2010 training compute grew in line with Moore's law, doubling roughly every 20 months. Since the advent of Deep Learning in the early 2010s, the scaling of training compute has accelerated, doubling approximately every 6 months. In late 2015, a new trend emerged as firms developed large-scale ML models with 10 to 100-fold larger requirements in training compute. Based on these observations we split the history of compute in ML into three eras: the Pre Deep Learning Era, the Deep Learning Era and the Large-Scale Era. Overall, our work highlights the fast-growing compute requirements for training advanced ML systems.

Motivation & Objective

  • Curate a dataset of milestone ML systems with training compute data.
  • Identify and characterize distinct eras of compute growth in ML.
  • Provide estimates of compute doubling times and discuss implications for hardware and research incentives.

Proposed method

  • Assemble a dataset of 123 milestone ML models with training compute annotations.
  • Fit log-linear models to training compute (FLOPs) over time to estimate doubling times.
  • Segment the history into Pre Deep Learning, Deep Learning, and Large-Scale eras and compare slopes and fit quality.
  • Cross-validate results with appendices addressing alternative interpretations and domain differences.

Experimental results

Research questions

  • RQ1What are the growth rates (doubling times) of training compute before and after the advent of Deep Learning?
  • RQ2Does a distinct Large-Scale era emerge around 2015-2016, and how does it differ from regular-scale compute growth?
  • RQ3How well do the identified trends fit the Milestone ML model data and what uncertainties exist in the estimates?

Key findings

  • Pre Deep Learning Era shows compute increasing roughly with Moore’s law, doubling about every 21 months (1952–2010).
  • Deep Learning Era speeds up compute growth to about a 5–6 month doubling time (2010–2022).
  • Large-Scale Era emerges around 2015–2016, with models exceeding prior trends and doubling approximately every 10 months (late 2015–2022).
  • Overall, three eras capture discontinuities in compute trends and highlight rising compute demands for advanced ML systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.