Shot saturation and evidence limits in quantum-AI hardware benchmarks
Do more quantum measurements produce a better answer? A reanalysis of hardware records tracks the error that remains as measurement budgets grow.
When more measurements stop helping.
Archived overlap measurements
More shots. Almost the same error.
Repeating a measurement reduces sampling variation. A persistent offset can still dominate the result.
Shots per batch · Doubling scale
Eight archived signed-overlap submissions; shot settings reuse nested data. The decomposition describes these records and does not identify a physical noise mechanism. Source: “Overlap results” in the Discover Quantum Science manuscript.
What the study shows.
More shots leave most error intact.
Across eight overlap submissions, an eightfold increase in shots changes reconstructed RMSE from 0.1335 to 0.1292. At 2048 shots, the squared mean-offset component accounts for 98.2% of the reconstructed error.
Missing entries are not zeros.
Only 18 of 45 intended off-diagonal kernel entries were measured. Zero-filling the remainder changes the mean from 0.2220 over observed entries to 0.0888 over all entries.
Compare the same quantity.
QAOA edge-cut fractions and optimum-normalized ratios use different denominators. Missing histograms and inconsistent source labels limit the comparisons that the archive supports.
Scope & details.
This is a retrospective analysis of archived summaries, with no new QPU execution. The error decomposition does not identify a physical noise mechanism, establish a general provider ranking, or demonstrate quantum advantage.
Read the abstract
Quantum benchmarks for artificial intelligence require measurement budgets, target observables, and execution records to agree. We reanalyze archived signed-overlap, kernel, and quantum approximate optimization algorithm summaries from a cloud-hardware campaign, retaining incomplete runs and distinguishing measured entries from imposed values. For eight signed-overlap submissions, increasing shots per batch from 256 to 2048 changes reconstructed root-mean-square error from 0.1335 to 0.1292. At 2048 shots, squared offsets of the empirical batch means account for 98.2% of the reconstructed error. This decomposition is descriptive; it does not identify a physical noise mechanism. An incomplete kernel contains 18 measured off-diagonal entries out of 45; filling the others with zero changes its observed-entry mean from 0.2220 to 0.0888. A source-label conflict prevents a matched provider comparison. Across 43 Rigetti QAOA submissions, size-group mean edge-cut fractions range from 0.4782 to 0.5010, but heterogeneous parameters and absent raw histograms limit interpretation.
Citation
@unpublished{lex2026shots,
title={Shot saturation and evidence limits in quantum-AI hardware benchmarks},
author={Lex, Vikram},
year={2026},
note={Manuscript under review at Discover Quantum Science}
}