Skip to content
Methodology v1.0

How a verdict is issued, and what could flip it.

Everything here is designed so you do not have to trust us: criteria registered before the run, data sealed and hash-anchored, results you can recompute yourself.

A verdict is not binary: it has five layers

A verdict is not binary: it has five layers

The binary — did it win or not — is given away free by the literature. What is charged for is layers 2 to 5.

Layer 0 — Integrity

Do the claim’s artifacts run? runnable / partially / not runnable. Many claims die here.

Layer 1 — Verdict

Under a criterion pre-registered and sealed BEFORE anything runs. Never adjusted by looking at the result.

Layer 2 — Margin

Percentage gap to the champion baseline, time-to-solution ratio, number of seeds and interval — or their declared absence. If variance was not measured, we say it was not measured.

Layer 3 — Robustness

Sensitivity to size, seeds and hyperparameters, and the strength of the baseline: does it only win against weak ones?

Layer 4 — Crossover and triggers

Two or three observable events that would flip the verdict. It is what makes the report age well.

Layer 5 — Economic translation

The result in the client’s units. The client’s numbers come from the client or are marked as assumptions; we supply the arithmetic and the sources.

The evidence scale (EL0–EL5)

The evidence scale (EL0–EL5)

Adapted from GRADE, the evidence-review standard in medicine: a base level set by the design of the evidence, plus declared factors that raise or lower confidence — never hidden arithmetic, always the reason in writing.

levelnamewhat it requires
EL0Claim onlyA dated, sourced public claim. No raw data or code downloadable by hash.
EL1First-party artifactsRaw data and code published and hash-verifiable: anyone can recompute what the author did.
EL2Pre-registered & sealedCriterion registered BEFORE the run, run sealed, the result meets its own criterion, and a classical baseline declared on the same field. Still first-party.
EL3Independently reproducedA third party with no involvement or payment from the author reproduced the main result from the published artifacts. Today: empty. Nobody in the registry has it.
EL4Adversarially testedEL3 plus surviving a serious published refutation attempt, or being reproduced against the strongest known classical baseline, at parity.
EL5Verified in productionIndependent verification in a real deployment with economic consequence and a named external referee. Today, one: certified randomness.

The scale measures the strength of the evidence that the advantage is real. It does NOT measure survival over time: that is the status (surviving / contested / eroded / open / negative-selfpublished), which is orthogonal and shown alongside. A claim can be EL2 and eroded; another EL0 and surviving.

The v1.0 rubric is ready to seal and is NOT sealed yet. Until it is, we publish no claim’s level: scoring with a yardstick that can still be changed is exactly what this firm exists not to do.

juez-v1: budget parity

juez-v1: budget parity

The protocol that makes a comparison comparable. The classical and the quantum side get the same instance, the same compute budget and the same wall-clock, with fixed seeds and declared versions. The baseline is the champion of the class, tuned — not the first thing that compiles. Without declared parity a result is not evidence: it is a demo.

The Two-Layer Rule

The Two-Layer Rule

Layer A — the sealing infrastructure: mechanical, opinion-free. Anyone may pay for it, vendors included. It notarizes integrity and date, not truth. Layer B — judgment: verdicts, evidence levels, reports, Monitor, Registry, Library. Paid for only by the demand side; never by the party being evaluated, directly or indirectly.

Verify a seal yourself → · The full registry →

Version 1.0 · 1 Sep 2026. Methodology changes are published with dates; verdicts cite the version they were issued under.