Why does no one maintain a neutral quantum verification ledger?
Status as of: August 2026.
No economically independent body runs every quantum machine under the same rule — same instances, same compute budgets, re-runnable artifacts, verdicts revised as hardware changes. As of August 2026 that ledger does not exist, and its absence is not an oversight. Every actor in the ecosystem has a rational reason not to maintain it: vendors measure with the ruler their own architecture favors, academia is rewarded for novelty rather than maintenance, analysts bill the industry they cover, and public capital funds machines rather than arbitration. The three most serious partial attempts — DARPA's Quantum Benchmarking Initiative, Metriq, and QOBLIB — each cover a piece of the job. None covers all of it.
What exactly is missing?
Two facts, both already on the record. First: there is no published conversion between the metrics vendors report — Quantum Volume, algorithmic qubits, CLOPS, error per layered gate. A buyer comparing two machines is comparing two rulers; we mapped the four rulers and their records here. Second: the correction record runs on volunteers. Every major quantum claim later revised — IBM's kicked-Ising utility result, D-Wave's spin-glass simulation, the QML exponential speedups — was re-run by outside groups on their own time, using whatever public artifacts happened to exist (the reversal table, the dequantization record).
What is missing is the institution in between: a standing body, economically independent of every machine it measures, that runs the same instances under the same budgets on all of them, publishes re-runnable artifacts per verdict, and revises those verdicts as machines move. Call it a neutral verification ledger. Nobody maintains one.
Why each actor rationally passes
No conspiracy is required. Each actor's incentive points away from the job, and each is acting rationally within it — this section is structural analysis, not an accusation.
Vendors. A vendor's benchmark budget converts into the metric its architecture wins (the record of who defined, adopted and abandoned which metric). Funding a ledger where your machine can lose to a rival on a workload you did not pick is, from a fiduciary standpoint, a strange use of shareholder money. Every vendor reports honestly within its own definition; no vendor is paid to maintain the definition that compares them.
Academia. The academic unit of reward is the novel paper, not the maintained registry. The one large audit of the gap between policy and artifacts: of 204 computational papers published in Science under a mandatory data-and-code policy, 36% of the 180 contactable groups provided any materials and 26% of findings could be reproduced (Stodden, Seiler & Ma, PNAS 115, 2584, 2018 — we covered the mechanism here). A live ledger is pure maintenance — the activity the incentive system pays for least.
Analysts. Market-research firms sell reports and briefings, and the covered industry is also the customer base. That is a normal commercial structure, not a scandal — but it is not economic independence from the measured, which is the first requirement of an arbiter.
Governments. Public and private quantum capital overwhelmingly buys machines, not arbitration: of the $12.6B invested in quantum-technology startups in 2025, ~90% went to quantum computing companies, and the government share of startup funding fell from 33% (2024) to 3% (2025) as private capital moved in (McKinsey Quantum Technology Monitor 2026, Apr 2026 — the full capital split). The one public program that funds evaluation — DARPA's QBI, below — evaluates roadmaps toward 2033, not today's machines on user workloads.
Community. The initiatives that most resemble a ledger (Metriq, QOBLIB) run on foundation and consortium budgets, and take results as submissions rather than executing the machines themselves — coverage without arbitration.
Who is actually trying — and what each attempt does not do
Three initiatives are doing serious work adjacent to the gap. None is the ledger.
DARPA Quantum Benchmarking Initiative (QBI). The only arbiter-shaped effort with public money and independence from the measured. Since mid-2024 it has evaluated ~20 companies' concepts for utility-scale quantum computers — "utility-scale" defined as computational value exceeding cost. Eleven teams advanced to Stage B on 6 Nov 2025, with up to ~$15M each over ~1 year: IBM, IonQ, Quantinuum, QuEra, Xanadu, Atom Computing, Photonic, Diraq, Quantum Motion, Nord Quantique and Silicon Quantum Computing. At Stage C, government verification-and-validation teams test whether each machine can be built and operated as designed; two performers from the earlier US2QC pilot are already there. In March 2026 QBI opened a fresh Stage A call for architectures not yet evaluated (abstracts due 31 Jul 2026). Its managing director's framing: "it now seems likely that someone will build a utility-scale quantum computer by 2033, but it remains unclear exactly which team or teams might get across that finish line" (Micah Stoutimore, Mar 2026). What QBI is not: a ledger of today's machines. It evaluates vendor roadmaps and prototypes against a 2033 milestone; it does not run user workloads on every current QPU under one rule, and whether its Stage C evaluations will be published as public, comparable verdicts is not stated.
Metriq (Unitary Foundation). A community benchmarking platform, relaunched in 2026 around the metriq-gym suite with a composite score per platform (arXiv:2603.08680, 2026). It is the closest thing to a live public scoreboard — and its results are community-submitted via CLI rather than executed by an independent operator, and the platform currently labels itself beta: "data and visuals are still stabilizing" (metriq.info, accessed Aug 2026). Aggregation without arbitration.
QOBLIB (Quantum Optimization Working Group). The instance-and-baseline layer, done right: 10 problem classes, 1,200+ instances (500+ solved to proven optimality), transparent classical baselines, and a live leaderboard that has logged 2,000+ submitted results — Nature Computational Science paper and new website launched 12 Aug 2026 (IBM Quantum blog; we covered QOBLIB as the reproducibility standard). What it does not have: an executor. Submitters run their own hardware and report; the leaderboard displays best-known results, it does not re-run them.
QED-C application benchmarks. Open-source, vendor-neutral benchmark suites anyone can run (Lubinski et al., arXiv:2110.03137; IEEE TQE 2023). Code, not an institution: the consortium publishes tools; it does not operate a standing cross-vendor ledger.
Peer review. Reviews the text of a claim; does not re-run the experiment. The measured outcome of relying on it is the 36%/26% audit above.
| Initiative | What it does | What it does not do | Source, date |
|---|---|---|---|
| DARPA QBI | Public-budget V&V of vendor roadmaps to 2033; 11 teams in Stage B | Does not run user workloads on today's machines; public comparability of Stage C unstated | DARPA, 6 Nov 2025 / 11 Mar 2026 |
| Metriq | Live community scoreboard, metriq-gym suite, composite score | Results self-submitted; no independent executor; in beta | metriq.info; arXiv:2603.08680 (2026) |
| QOBLIB | 1,200+ instances, classical baselines, leaderboard with 2,000+ results | No arbiter executes; leaderboard lists submissions | Nature Comput. Sci. 2026; IBM blog, 12 Aug 2026 |
| QED-C suite | Open-source application benchmarks anyone can run | Tooling, not an institution; no standing ledger | arXiv:2110.03137 (2021; IEEE TQE 2023) |
| Peer review | Checks the text of claims | Does not re-run; 26% reproduced in the one large audit | PNAS 115, 2584 (2018) |
What would a neutral ledger require?
Four properties, each derived from how published claims have actually failed (weak baselines, missing artifacts):
Economic independence from the measured. The operator's budget cannot come from the vendors being ranked. QBI has this (public money); nothing else fully does.
Same instance, same budget, both sides. Quantum and classical run the identical problem under the identical compute budget. QOBLIB supplies the instances and baselines for one domain; no one supplies the execution.
Re-runnable artifacts per verdict. Data, code, seeds, instance and budget, published together per run. QOBLIB and Metriq both push in this direction — for submitted results.
Live, dated, revisable verdicts. Machines change monthly; an unmaintained ledger is a snapshot, and snapshots are what the correction-by-volunteers record is made of.
Each requirement has at least one partial owner. No initiative holds all four — that intersection is the gap.
Our own measured data
Rosetta's sealed-run protocol is a worked example of requirements 2 and 3 at small, declared scope: each run in our V-0012 / RQ series seals the instance, seeds, compute budget and result together, and the classical baseline remains unbeaten in every sealed run so far (the portfolio series). That is one operator, a handful of optimization and quantum-walk classes, and zero claim to be the neutral ledger — the scope statement is the point. A real ledger would need the same artifact discipline at ecosystem scale, operated by an institution with the independence properties above.
What we don't know
- Whether QBI's Stage C will produce public, comparable evaluations. DARPA describes government V&V of whether each machine can be built and operated as designed — not a published cross-vendor scoreboard.
- Whether the gap is transitory. A regulator, a standards body or a consortium could decide to fund arbitration; nothing structural forbids it. The incentive analysis above explains absence — it does not prove permanence.
- Whether QOBLIB's leaderboard becomes the de facto standard for optimization, and whether Metriq stabilizes out of beta into a maintained scoreboard.
- Whether anyone maintains such a ledger privately — a fund, a lab, a defense agency. The absence of a public ledger is not evidence that nobody keeps one.
- Our incentive matrix is structural inference from public behavior, not a measured survey. Nobody has polled vendors, labs and agencies on why they do not fund arbitration.
Sources
- DARPA — QBI Stage B selection (6 Nov 2025)
- DARPA — QBI expands, new Stage A call (11 Mar 2026)
- The Quantum Insider — QBI expansion coverage (11 Mar 2026)
- PostQuantum — QBI Stage B analysis: funding and stages (Nov 2025)
- IBM Quantum blog — QOBLIB: tracking progress in quantum optimization (12 Aug 2026)
- QOBLIB — The Quantum Optimization Benchmarking Library (Nature Computational Science, 2026)
- Metriq platform (Unitary Foundation, accessed Aug 2026)
- Metriq — A Collaborative Platform for Benchmarking Quantum Computers (arXiv:2603.08680, 2026)
- Stodden, Seiler & Ma — reproducibility audit of Science papers (PNAS 115, 2584, 2018)
- McKinsey Quantum Technology Monitor 2026 (Apr 2026)
- Lubinski et al. — QED-C application-oriented benchmarks (arXiv:2110.03137; IEEE TQE 2023)
Rosetta Q publishes verdicts with reproducible raw data. This is educational content, not a product claim.