>_ Skip to main content
Menu
Search
Quantum Technology

QUOPS Benchmark Finds 100,000-Fold Quantum Performance Gap

For years, qubit counts have been the primary metric in quantum computing. However, a new benchmark called QUOPS argues that this number reveals very little about a machine’s ability to perform real-world tasks. The tool indicates that current hardware is around five orders of magnitude, or 100,000 times less powerful than what’s needed for scientifically useful computation.

QUOPS, which stands for Quantum Universal Operations Performance System, was developed by Sandia National Laboratories with contributions from researchers at Quantinuum and NVIDIA. The system is described in a paper posted to the arXiv preprint server on September 15. The paper has not yet undergone peer review.

How QUOPS Measures Performance

Most vendor specifications highlight qubit count, gate accuracy, and gate speed. Though these figures describe individual components, they don’t illustrate how an entire machine performs during large calculations, where errors accumulate and components interact.

QUOPS, instead, assesses the computer as a single system. Researchers input randomized circuits of varying width and depth, then verify if the output meets a specified accuracy threshold. “Width” refers to the number of qubits used, while “depth” indicates the number of sequential operations. A wide, shallow circuit distributes work across many qubits over few steps, whereas a narrow, deep circuit uses fewer qubits for a longer duration.

By running numerous circuit shapes, the team can map a “capability region,” and show which width and depth combinations a particular machine can handle above the threshold. This region is then condensed into two numbers:

  1. Q: Represents the largest qualifying circuit within a size range designed to mimic practical workloads. Circuit size here is defined as twice the width multiplied by the depth.
  2. Omega (Ω): The effective number of quantum operations completed per second at that specific point.

These two figures highlight different aspects of performance. A high Q with a low Ω describes a machine that can complete large circuits slowly. Conversely, a high Ω with a small Q indicates a fast machine limited to short calculations. The authors specifically chose two metrics to prevent a single number from obscuring this crucial tension.

Performance Across Three Architectures

The researchers applied QUOPS to physical qubits from Quantinuum, Google, and IBM. This cross-platform testing is significant because these companies employ different hardware physics.

Google’s Willow and IBM’s Boston systems utilize superconducting qubits, which execute gates rapidly. However, each qubit connects to only a few neighbors. Operations between distant qubits require additional routing, which prolongs circuits and increases error potential. Quantinuum’s Helios, on the other hand, uses trapped ions in a quantum charge-coupled device design. Its gates are slower, but it offers effective all-to-all connectivity, allowing any pair of qubits to interact without the routing overhead.

The results reflected these trade-offs. According to Quantinuum’s blog post, Willow and Boston achieved higher QUOPS rates but exhibited smaller capability regions and lower Q scores. Helios, in contrast, managed larger circuits and a higher Q score, albeit at a slower rate. Given Quantinuum’s co-authorship and Helios being its product, it’s important to consider this framing.

No single architecture emerged as universally superior. Systems achieved performance through distinct combinations of accuracy, speed, and connectivity, with the optimal mix depending on the specific workload. QUOPS also accounted for error mitigation, a technique involving running circuits multiple times and post-processing the data to derive a cleaner result. While mitigation can expand circuit capabilities, the extra sampling reduces the effective operations per second. The paper’s rate analysis quantified this trade-off.

Testing Logical Qubits

The team further applied QUOPS to a small fault-tolerant setup on Quantinuum’s Helios-1 system. This involved using up to eight logical qubits, each encoded with a seven-qubit error-correcting code. Logical qubits are formed by groups of physical qubits and can detect and correct a limited number of errors.

This was not a demonstration of a fully functional fault-tolerant computer but rather a small experiment to illustrate that the same benchmark can measure both current physical-qubit machines and nascent logical-qubit systems. 

This continuity is important because as developers adopt different error-correction codes, raw physical-qubit counts become less comparable; one company might require significantly more physical qubits than another to build a logical qubit of similar quality. QUOPS bypasses this issue by measuring the circuits a system can actually execute.

The 100,000-Fold Gap: An Estimate, Not a Final Word

The headline-grabbing finding is that current systems are approximately five orders of magnitude less powerful than the computational capability estimated for several recognized scientific problems. This translates to a 100,000-fold performance increase needed to reach the target range analyzed by the researchers.

This estimate supports the argument for fault-tolerant quantum computing, where error correction protects long calculations from accumulating failures. Though today’s machines perform increasingly sophisticated experiments each year, noise still limits the length and size of reliable computations. This study does not report a quantum advantage; it measures capability against a projected target, not a commercially solved problem.

The limitations of this assessment warrant attention. QUOPS uses randomized circuits as a stress test, not full chemistry, materials, or optimization applications, so passing it doesn’t guarantee a machine will deliver useful results on specific problems. The summary score relies on chosen inputs, such as the success threshold and the range of circuit shapes deemed relevant. 

The logical-qubit test involved eight encoded qubits and a simple code, and projections to future systems are based on assumptions about error rates and scaling that have not been demonstrated at application scale. Furthermore, access remains an issue: vendors control most advanced processors, and results can vary with compilation and calibration. Independent, repeatable comparisons will require published test rules and access to raw data.

Implications for Buyers

Quantinuum states that QUOPS is not intended to replace component measurements or application-specific tests. The company encourages hardware developers to report Q and Ω alongside existing specifications and urges agencies and buyers to incorporate QUOPS thresholds into procurement requests. This request aligns with Quantinuum’s interests, given Helios’s performance on the Q score.

If the benchmark gains adoption and independent verification, it could provide computing centers and government buyers with a clearer method for comparing systems. A buyer seeking a machine capable of a trillion reliable operations could set a system-level target without committing to a specific qubit technology or error-correction scheme. For now, the paper is a preprint, and peer review is the next step. For full details, refer to the arXiv paper, and consider the 100,000-fold figure an estimate that will evolve with improved algorithms, codes, and hardware.