Why does a weak classical baseline ruin a quantum benchmark?
A quantum benchmark is a subtraction. You measure a quantum method's cost, you measure a classical method's cost on the same problem, and the claim is the gap between them. That means the number you publish is only as honest as the classical side of the subtraction. Pick a weak opponent and you can manufacture any advantage you like — no better hardware required.
This is the most common way a quantum benchmark lies, and it is almost never malicious. It is a strong quantum team that is not also a strong classical team, running the classical baseline at default settings, or reaching for the wrong classical algorithm, and reporting the gap in good faith. The result reads as physics. It is arithmetic.
What exactly is a "baseline"?
The baseline is the classical method you race against: the specific solver, its tuning, the hardware it runs on, and the wall-clock or energy budget it gets. "Classical computers need X" is not a baseline — it is a sentence. A baseline is Gurobi 11 with these parameters, on this CPU, given the same 60 seconds. The strength of that opponent sets the entire meaning of the result. A benchmark against a strong baseline is a measurement. A benchmark against a strawman is a press release.
Why a weak baseline moves the crossover point for free
Ties directly to the crossover point: the problem size N* where the quantum method finally becomes cheaper than the best classical one. N* is where a speedup stops being academic and starts being real. But N* is defined against a specific classical curve. Weaken that curve — raise its cost — and its intersection with the quantum curve slides left. You have not moved the physics; you have moved the opponent, and the crossover moved with it.
Read the gold curve and the teal curve as the same classical problem solved two ways. The weak baseline (gold, dashed) blows up fast and crosses the quantum curve at a small "fake N*" — quantum appears to win. The strong baseline (teal) stays cheaper than quantum at every size in range: no crossover exists at all. Nothing about the quantum device changed between the two pictures. Only the opponent did.
The record: claims that evaporated when a strong baseline showed up
This is not hypothetical. The most-cited "beyond-classical" results of the last three years were, in each case, matched by a classical method that arrived after the headline — usually tensor networks, often on hardware a scientist already owns.
| Claim | What was asserted | Strong classical baseline that matched it | Source |
|---|---|---|---|
| IBM 127-qubit kicked Ising, 2023 | expectation values "beyond brute force" | tensor-network methods reproduced them on modest compute, 2023–24 | Kim et al., Nature 618, 500 (2023); Tindall et al., PRX Quantum 5, 010308 (2024) |
| D-Wave Advantage2 spin-glass quench, Mar 2025 | quench dynamics past classical reach | 2D/3D tensor networks reproduced key regimes on a laptop, 2025–26 (disputed) | Tindall et al., arXiv:2503.05693 (2025) |
| QML exponential speedups, ~2018 | exponential gap over classical | dequantized classical algorithms with comparable scaling | Tang, arXiv:1807.04271 (2018) |
The pattern is not that quantum hardware is fake. IBM's expectation values were real; D-Wave's device did what it did; the QML circuits ran. The pattern is that the classical baseline in the original comparison was not the strongest one available, and the gap shrank — sometimes to zero — once it was. Read each as a measurement corrected, not a fraud exposed. The D-Wave case in particular is still live: D-Wave disputes that the tensor-network rebuttal covers the full claimed regime. That dispute is itself the point — the argument is now about whose classical baseline is strong enough, which is exactly the right argument to be having.
How to tell if a baseline is strong enough
Before you believe any advantage number, interrogate the classical side, not the quantum one:
- Best-known, not convenient. Is the baseline the state-of-the-art classical solver for this exact problem class (Gurobi/CP-SAT for optimization, tensor networks for these spin systems, DiffDock-class for docking), or the first library that imported cleanly?
- Tuned, not default. Was the classical solver actually configured by someone who wanted it to win, or run at stock settings?
- Same instance, same budget. Identical problem instances, identical wall-clock or energy budget, identical hardware class on both sides?
- Reported, not narrated. Is the classical result a measured number with a seed and a config you can rerun — or a sentence ("classical would take years") with no artifact behind it?
- Reproducible. Can an outsider rerun both sides and get the same gap?
A "no" to any of these is enough to withhold belief. Not to conclude the quantum side is wrong — to conclude the benchmark hasn't earned a verdict yet.
What we ran ourselves
We hold our own verdicts to rule 1, and it costs us positive results. In verdict V-0012 (RQ-0012), portfolio optimization, QAOA at depth p=2 against CP-SAT on 20 sealed runs at n=12/16/20: CP-SAT proved the optimum 20 out of 20, while QAOA landed 25–48% away with no size trend. If we had graded QAOA against a weak baseline — random sampling, or an untuned local search — QAOA would have "won," and that win would have been an artifact of a bad opponent, not a property of the algorithm. We use a strong baseline on purpose, and against it there is no crossover. Note the asymmetry that makes this honest: our quantum runs are noiseless simulations, which favor the quantum side, and it still loses to a well-tuned classical solver.
What we don't know
We have not independently re-run the tensor-network rebuttals of the IBM or D-Wave results; we are reporting the published record and the ongoing dispute, not adjudicating it. "Strongest known classical baseline" is a moving target — a method strong enough today can be beaten by a better classical algorithm tomorrow, which is exactly why crossover claims decay. We cannot prove a negative: the absence of a classical method that beats a given quantum result is never final, only current. And a strong baseline is necessary, not sufficient — same instance, same budget, and reproducibility all have to hold too.
Rosetta Q publishes verdicts with reproducible raw data. This is educational content on measurement methodology, not a product claim. A benchmark is only as strong as the classical baseline it beats; we hold ours to the strongest one we can find, and we publish the negatives.