Skip to main content
QUICK REVIEW

[Paper Review] Benchmarking TinyML Systems: Challenges and Direction

Colby Banbury, Vijay Janapa Reddi|arXiv (Cornell University)|Mar 10, 2020
Video Analysis and SummarizationComputer Science24 references197 citations
TL;DR

The paper discusses the need for a fair hardware benchmark for TinyML, outlines key challenges, and proposes a four-benchmark TinyMLPerf-like suite with open/closed divisions, using four diverse use cases, datasets, and models.

ABSTRACT

Recent advancements in ultra-low-power machine learning (TinyML) hardware promises to unlock an entirely new class of smart applications. However, continued progress is limited by the lack of a widely accepted benchmark for these systems. Benchmarking allows us to measure and thereby systematically compare, evaluate, and improve the performance of systems and is therefore fundamental to a field reaching maturity. In this position paper, we present the current landscape of TinyML and discuss the challenges and direction towards developing a fair and useful hardware benchmark for TinyML workloads. Furthermore, we present our four benchmarks and discuss our selection methodology. Our viewpoints reflect the collective thoughts of the TinyMLPerf working group that is comprised of over 30 organizations.

Motivation & Objective

  • Motivate the need for a fair, comparable TinyML hardware benchmark to accelerate progress.
  • Survey the TinyML landscape across use cases, models, and datasets to identify benchmarking gaps.
  • Identify fundamental challenges (power, memory, hardware/software heterogeneity) in TinyML benchmarking.
  • Propose a concrete path forward with four benchmark use cases, datasets, and reference models.

Proposed method

  • Analyze the current TinyML landscape and benchmarking efforts.
  • Identify core challenges that complicate fair benchmarking (power, memory, hardware and software heterogeneity).
  • Propose principled guidelines for a TinyML benchmark suite and define four target use cases.
  • Select open datasets and reference models to ground the closed division benchmarks.
  • Define a measurement framework emphasizing latency with optional energy metrics and divisions for comparability.

Experimental results

Research questions

  • RQ1What are the principal challenges in creating a fair and useful TinyML hardware benchmark?
  • RQ2How should a TinyML benchmark be structured to balance comparability, openness, and representativeness?
  • RQ3Which use cases, datasets, and models best cover the TinyML landscape for benchmarking purposes?
  • RQ4What metrics and divisions (open vs. closed) best enable fair evaluation across heterogeneous TinyML hardware?

Key findings

  • TinyML benchmarking faces four main challenges: low power measurement, extreme memory constraints, hardware heterogeneity, and software deployment diversity.
  • A TinyML benchmark should adopt open and closed divisions to balance strict comparability with inclusivity and innovation.
  • Four use cases were selected (audio wake words, visual wake words, image classification, anomaly detection) to cover diverse input types and model families.
  • Open division results must keep accuracy within a threshold of the closed-division reference model.
  • Metrics focus on inference latency with an option to measure energy consumption.
  • A rapid, minimum-viable benchmarking set is prioritized with plans for iterative improvement and community involvement.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.