Skip to content
Under reviewDiscover Quantum Science

Shot saturation and evidence limits in quantum-AI hardware benchmarks

Do more quantum measurements produce a better answer? A reanalysis of hardware records tracks the error that remains as measurement budgets grow.

Inside the study

When more measurements stop helping.

Archived overlap measurements

More shots. Almost the same error.

Repeating a measurement reduces sampling variation. A persistent offset can still dominate the result.

256 shots2,048 shots
0.1292Reconstructed RMSE
Total measured errorRMSE · Scale 0–0.16
2565121,0242,048

Shots per batch · Doubling scale

98.2%of reconstructed mean squared error at 2,048 shots is the squared mean-offset component.

Eight archived signed-overlap submissions; shot settings reuse nested data. The decomposition describes these records and does not identify a physical noise mechanism. Source: “Overlap results” in the Discover Quantum Science manuscript.

The findings

What the study shows.

More shots leave most error intact.

Across eight overlap submissions, an eightfold increase in shots changes reconstructed RMSE from 0.1335 to 0.1292. At 2048 shots, the squared mean-offset component accounts for 98.2% of the reconstructed error.

Missing entries are not zeros.

Only 18 of 45 intended off-diagonal kernel entries were measured. Zero-filling the remainder changes the mean from 0.2220 over observed entries to 0.0888 over all entries.

Compare the same quantity.

QAOA edge-cut fractions and optimum-normalized ratios use different denominators. Missing histograms and inconsistent source labels limit the comparisons that the archive supports.

A closer look

Scope & details.

This is a retrospective analysis of archived summaries, with no new QPU execution. The error decomposition does not identify a physical noise mechanism, establish a general provider ranking, or demonstrate quantum advantage.

Read the abstract

Quantum benchmarks for artificial intelligence require measurement budgets, target observables, and execution records to agree. We reanalyze archived signed-overlap, kernel, and quantum approximate optimization algorithm summaries from a cloud-hardware campaign, retaining incomplete runs and distinguishing measured entries from imposed values. For eight signed-overlap submissions, increasing shots per batch from 256 to 2048 changes reconstructed root-mean-square error from 0.1335 to 0.1292. At 2048 shots, squared offsets of the empirical batch means account for 98.2% of the reconstructed error. This decomposition is descriptive; it does not identify a physical noise mechanism. An incomplete kernel contains 18 measured off-diagonal entries out of 45; filling the others with zero changes its observed-entry mean from 0.2220 to 0.0888. A source-label conflict prevents a matched provider comparison. Across 43 Rigetti QAOA submissions, size-group mean edge-cut fractions range from 0.4782 to 0.5010, but heterogeneous parameters and absent raw histograms limit interpretation.

Citation
@unpublished{lex2026shots,
  title={Shot saturation and evidence limits in quantum-AI hardware benchmarks},
  author={Lex, Vikram},
  year={2026},
  note={Manuscript under review at Discover Quantum Science}
}