Pillar D · State as of 2026-07-23

Runs 003–008: variance dissolves the trend (and that's the system working)

We ran 3 new seeds per size. At n=12 the QAOA gap spans 23.4%–53.7% (mean ≈40.9%); at n=16 it spans 20.7%–42.3% (mean ≈28.8%). The ranges overlap: run 002's apparent 'gap shrinks with size' does not survive measurement. Classical CP-SAT hit the proven optimum in every one of the 8 runs to date. Verdict, all runs: not yet. Six sealed archives, EXP-0012-003 through -008.
→ Leer en español
State as of: 2026-07-23

Why this batch existed

Run 002 showed a smaller gap at n=16 than at n=12 (20.7% vs 42.8%) and we explicitly refused to call it a trend: one seed per size, unmeasured variance. This batch measures that variance — three fresh seeds (43, 44, 45) per size, protocol byte-for-byte frozen: same budgets, same depth, same solver, exact-optimum referee.

The result

QAOA p=2 · gap to exact optimum · 4 seeds per size seeds 42–45 · equal 120 s budgets · CP-SAT: proven optimum in all 8 runs 0% 60% n=12 n=16 max 53.7% min 23.4% mean 40.9% max 42.3% min 20.7% mean 28.8% teal dots = individual seeds · gold tick = mean ranges overlap → size effect unresolved at n=4 seeds
Measured results, sealed archives EXP-0012-001 … -008.
n=12  gaps: 42.8 · 53.7 · 43.7 · 23.4   mean 40.9   range 30.3
n=16  gaps: 20.7 · 42.3 · 21.3 · 31.0   mean 28.8   range 21.6

What the batch settles — and what it kills

It kills the run-002 story. “The gap shrinks as instances grow” looked plausible with one seed per size. With four, the ranges overlap by 19 points: seed 45 at n=12 (23.4%) beats seed 43 at n=16 (42.3%). The means still differ (40.9 vs 28.8), but with n=4 seeds and this spread, that difference is not yet distinguishable from noise. We said two points define any story you want; here is the proof, at our own expense.

It settles that variance is first-order. QAOA’s seed-to-seed spread (~20–30 points of gap) is the same magnitude as any size effect we could hope to measure. Every future claim on this curve must carry error bars or it is decoration.

And one thing stayed constant across all 8 runs: CP-SAT found the provably exact optimum every single time, in ≈0.1–1.1 s. The classical bar is not moving. That is what a strong baseline looks like.

What we don’t know

Whether the mean gap truly decreases with size — answering that needs more seeds per size (next tanda: 10+) and larger n. Whether deeper circuits (p=4) tighten the spread or just shift it. And how any of this behaves under hardware noise. All measurable; all on the ladder.

Measured content from sealed archives EXP-0012-001 … -008. Rosetta Quantum publishes verdicts with reproducible raw data — including the ones that dissolve our own prior observations.

Sources:
· Sealed archives EXP-0012-003 … -008, RosettaQuantum/evidence", url: "https://github.com/RosettaQuantum/evidence/tree/main/runs/2026/07" }