Skip to main content
QUICK REVIEW

[Paper Review] Defining Standard Strategies for Quantum Benchmarks

Mirko Amico, Helena Zhang|arXiv (Cornell University)|Mar 3, 2023
Quantum Computing Algorithms and Architecture10 citations
TL;DR

The paper defines universal benchmark criteria for quantum hardware, distinguishes benchmarks from diagnostics, and analyzes Quantum Volume, CLOPS, mirror circuits, and application benchmarks, including how optimizations and error mitigation should be reported.

ABSTRACT

As quantum computers grow in size and scope, a question of great importance is how best to benchmark performance. Here we define a set of characteristics that any benchmark should follow -- randomized, well-defined, holistic, device independent -- and make a distinction between benchmarks and diagnostics. We use Quantum Volume (QV) [1] as an example case for clear rules in benchmarking, illustrating the implications for using different success statistics, as in Ref. [2]. We discuss the issue of benchmark optimizations, detail when those optimizations are appropriate, and how they should be reported. Reporting the use of quantum error mitigation techniques is especially critical for interpreting benchmarking results, as their ability to yield highly accurate observables comes with exponential overhead, which is often omitted in performance evaluations. Finally, we use application-oriented and mirror benchmarking techniques to demonstrate some of the highlighted optimization principles, and introduce a scalable mirror quantum volume benchmark. We elucidate the importance of simple optimizations for improving benchmarking results, and note that such omissions can make a critical difference in comparisons. For example, when running mirror randomized benchmarking, we observe a reduction in error per qubit from 2% to 1% on a 26-qubit circuit with the inclusion of dynamic decoupling.

Motivation & Objective

  • Propose a set of benchmark postulates (randomized, well-defined, holistic, platform independent) and differentiate benchmarks from diagnostics.
  • Illustrate the criteria using existing benchmarks (Quantum Volume, CLOPS) and mirror circuits to discuss scalability and applicability.
  • Clarify how optimizations and error mitigation should be reported in benchmarking to avoid unfair comparisons.
  • Exhibit application-oriented and mirror benchmarking approaches to demonstrate the practical impact of the proposed rules.

Proposed method

  • Define benchmark postulates: randomization, well-defined procedures, holistic coverage, and platform independence.
  • Explain Quantum Volume with its depth- and circuit-based success criteria and discuss statistical passing rules.
  • Present CLOPS and its formula for throughput of QV layers per second.
  • Introduce mirror circuits as scalable benchmarks and describe their construction and trade-offs.
  • Discuss application-oriented benchmark suites and how they fit into the benchmark vs diagnostic framework.
  • Provide experimental illustrations showing the impact of optimization and mitigation techniques on benchmark outcomes.

Experimental results

Research questions

  • RQ1What constitutes a faithful, fair benchmark across diverse quantum hardware platforms?
  • RQ2How do different benchmark families (QV, CLOPS, mirror benchmarks, application suites) compare in scope, scalability, and noise sensitivity?
  • RQ3What optimizations are appropriate in benchmarking, and how should their overhead be reported?
  • RQ4How do error mitigation and suppression techniques affect benchmark results and comparability?

Key findings

  • Benchmarks should be randomized, well-defined, holistic, and platform independent to enable fair cross-platform comparisons.
  • A clear distinction between benchmarks and diagnostics helps avoid overgeneralizing performance from circuit-specific tests.
  • Mirror circuits offer a scalable, fast proxy for benchmarking large quantum systems while highlighting certain error sensitivities.
  • Optimization techniques can improve benchmark outcomes but must be transparently reported with their overhead and limitations.
  • Error mitigation can dramatically improve observed quality but often incurs exponential classical or quantum overhead, influencing comparability across devices.
  • Application-oriented benchmarks can illustrate practical performance but should be interpreted cautiously due to circuit-structure biases and potential optimizations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.