Pillar D · State as of 2026-09-04

Why can a detector look good just by being close to the active site?

In an internal audit dated September 2, 2026, Rosetta Q measured that its own de novo allosteric-site detector fires at the 11th percentile of distance to the active site while real allosteric sites sit at the 28th, and that once the statistical null is matched for that distance the project's headline p-value falls from 5.4e-6 to 0.0138. The result still crosses 0.05, but removing a single protein takes it to 0.057. Rosetta Q is publishing the correction because a control this cheap is one any reviewer can run in an afternoon.
→ Leer en español
State as of: 2026-09-04
ROSETTA QQUANTUM VERIFICATION LEDGERPILLAR D · RUN REPORTSTATUS AS OF: 2026-09-04Why can a detector lookgood just by being closeto the active site?11 / 28distance percentile to theactive site — where we point(11) vs. real sites (28)11OURS28REAL med56REAL p75rosettaquantum.com/blograw, reproducible data · sources below

What did Rosetta Q measure?

That its own detector points at the wrong shell. In an internal run sealed on 2026-09-02, the de novo allosteric-site detector of Rosetta Q was audited against a control it had not been using: distance to the active site. The pockets it predicts sit at the 11th percentile of distance to the active site, measured inside each protein; the real allosteric sites of our validated set of 86 proteins sit at the 28th (medians; Mann-Whitney, candidates closer than real sites, p < 1e-5, measured 2026-08-31). In Ångströms: 9.0 median for the 15 predicted pockets, 12.8 for the real ones. Only about one in ten real sites is as close to the active site as our median candidate.

Distance is legitimate signal, and that is the part worth defending. Real allosteric sites are closer to the active site than a random distal residue would be, so a model is right to use the feature. The failure is over-exploitation: the detector fires into a shell where roughly one in six real sites lives, and the scoring did not notice.

DISTANCE TO ACTIVE SITE · PERCENTILE PER PROTEINWHERE THE DETECTOR FIRESmedian 11 · IQR 7–15 · n=15WHERE REAL ALLOSTERIC SITES LIVEmedian 28 · IQR 18–56 · n=860%25%50%75%100%medians: 9.0 Å predicted vs 12.8 Å realMann–Whitney p < 1e-5 · Rosetta Q, 2026-08-31

Why does distance to the active site fool a detector?

Because the null it is scored against ignores distance. Our confidence number is the fraction of random contiguous pockets that score at least as high as the predicted one. When those random pockets are drawn from all distal residues, they land, on average, much farther from the active site than the real site does. A detector that has learned nothing except "closer is better" then beats that null without having learned anything about allostery.

The internal control makes the size of the gap concrete. For EGFR, the real allosteric site sits 8.3 Å from the active site while the historical null drew its pockets at a median of 16.5 Å. The comparison was never fair. This is the same failure Rosetta Q has written about in quantum benchmarking: a weak classical baseline ruins a benchmark because the number you publish is a statement about your control, not about your method. A weak null is that same sin in statistics.

What is a distance-matched null?

A null that draws its random pockets from the same distance shell as the thing being tested. Rosetta Q rebuilt it in a new module and changed exactly one thing: the seeds of the null pockets are drawn only among distal residues whose minimum distance to the active site is within ±2 Å of the median distance of that protein's real allosteric site. Same 10 features, same leave-one-protein-out protocol, same statistic, same 2,000 permutations, same Fisher combination. The question changes from "does this score high anywhere?" to "does this score high for how far away it is?"

The control does not punish everyone equally, which is how we know it is doing its job. Under the matched null EGFR degrades from 0.0065 to 0.049 (its site is in the 8th percentile of distance) while PAK1 improves from 0.0175 to 0.0005 (29th percentile, genuinely distal). It punishes exactly the sites that only stood out by being close.

What happened to the headline number?

It survived, and it fell by roughly three orders of magnitude. All 86 proteins stayed evaluable under the matched null (matched pools of 21 to 164 seeds, median 59), so both columns describe the same set.

Condition Fisher p (N=86) Proteins under p<0.05 Source
Historical null (contiguous, unmatched), seed A 5.4e-6 8 / 86 Rosetta Q run report, sealed 2026-09-02
Historical null, seed B 6.6e-6 8 / 86 Rosetta Q run report, sealed 2026-09-02
Matched null, ±2 Å band, seed A 0.0138 5 / 86 Rosetta Q run report, sealed 2026-09-02
Matched null, ±2 Å band, seed B 0.0107 5 / 86 Rosetta Q run report, sealed 2026-09-02
Matched null, ±1 Å band 0.0011 6 / 86 Rosetta Q run report, sealed 2026-09-02
Matched null, ±3 Å band 0.0076 5 / 86 Rosetta Q run report, sealed 2026-09-02

The same audit corrected a number we had been citing wrong. The internal ledger stored the combined p rounded to four decimals, so it read 0.0000, and an internal one-pager had turned that into "p is approximately 0, below the 1e-6 numerical floor". There was no floor: the real historical value at N=86 is 5.4e-6 — significant, but neither zero nor capped. Rosetta Q cites it as 5.4e-6 from now on, and the wrong version stays on the record rather than being quietly overwritten.

TWO NULLS, ONE METRIC · N=86HISTORICAL NULLno distance controlFisher p = 5.4e-6proteins under p<0.058 / 86drop the 5 strongeststill 0.004what it askshigh anywhere?MATCHED NULL ±2 Åsame distance shellFisher p = 0.0138proteins under p<0.055 / 86 vs 4.3drop 1 protein0.057 — no longerwhat it askshigh for how far?same data, same model, same 86 proteins — only the control changed

How fragile is what survives?

One protein fragile. The aggregate crossing rests on two cases, PAK1 and SLC6A4, both sitting at the resolution floor of the matched null (p = 1/2001 = 0.0005).

Proteins removed (strongest first) Matched-null Fisher p Historical-null Fisher p Crosses 0.05? Source
none 0.0138 5.4e-6 yes Rosetta Q run report, 2026-09-02
1 (PAK1) 0.057 5.4e-5 no Rosetta Q run report, 2026-09-02
2 (PAK1, SLC6A4) 0.176 2.2e-4 no Rosetta Q run report, 2026-09-02
3 0.266 7.4e-4 no Rosetta Q run report, 2026-09-02
5 0.435 3.7e-3 no Rosetta Q run report, 2026-09-02

There is a second reading that matters more than the aggregate. Under the matched null, 5 of 86 proteins come in below p = 0.05 against 4.3 expected by chance at that threshold: exact binomial p = 0.43. Protein by protein, the detector does not hit more often than chance. What is left is a weak aggregate effect concentrated in two cases, and the honest way to say it is exactly that.

MATCHED-NULL FISHER p vs k REMOVED (LOG)0.0138k=00.057k=10.176k=20.266k=30.435k=5threshold 0.05historical null at k=5: 0.004 · matched null breaks at k=1

What did the same control do to the rest of the work?

It removed 9 of 15. Applied on 2026-08-31 to the forward commitments in the discovery queue — candidate pockets flagged in proteins with no known allosteric site — the distance-matched null took 9 of the 15 below significance, including the candidate that had been the strongest in the project. Those commitments are dated bets, not findings, and they get their own report; the point here is that the same control cut in the same direction everywhere it was applied.

The bias also reproduced on fresh targets that had nothing to do with the original batch. Of five new targets opened on 2026-09-02, the two that passed the first filter sat at the 9th and 7th percentile of distance to the active site. Rosetta Q closed both. The periphery bias is not a property of one batch; it is a property of the detector, whose blind top-1 hit rate is 32.6% (28/86), measured 2026-08-28.

Is this failure mode specific to Rosetta Q?

No, and the fix has a long history in other fields. The general shape is that a model learns a property of how the negative examples were assembled rather than the biology, and the benchmark rewards it.

Field The bias The fix Source
Structure-based virtual screening Deep models scored well on DUD-E by learning ligand-side artifacts of how decoys were chosen, not protein-ligand interaction Property-matched decoys, and testing whether the model still works with the protein hidden Chen et al., PLOS ONE 14(8), 2019
Benchmark design in cheminformatics Decoys that differ from actives in trivial physical properties make retrieval look easy Decoys matched on physical properties but topologically dissimilar (the DUD-E construction itself) Mysinger et al., J. Med. Chem. 55, 6582 (2012)
Machine-learning-based science, 17 fields reviewed Leakage inflates reported performance across hundreds of papers; the model sees, indirectly, what it is meant to predict Model info sheets and explicit auditing of what the evaluation set shares with training Kapoor & Narayanan, Patterns 4, 100804 (2023)
Genomics Distributional differences between positive and negative sets are a documented, recurring pitfall Match negatives on the confounding property; report the matched result Whalen et al., Nat. Rev. Genet. 23, 169 (2022)
Allosteric-site detection (this report) Predicted pockets sit at distance percentile 11, real sites at 28; the unmatched null does not control for it Seeds drawn from the same ±2 Å distance shell as the real site Rosetta Q run report, sealed 2026-09-02

Nothing here is an accusation of anyone, including us. An unmatched null is the default construction, not misconduct, and it is invisible until somebody measures the property it fails to control. What is worth judging is what happens after the measurement exists.

What does this change in how Rosetta Q reports?

Four things, all of them dated. The matched null becomes the primary metric of the project. The historical value is reported beside it, never in its place, because the two are not comparable and the matched one is stricter. The corrected number, 5.4e-6, replaces "approximately 0" everywhere it was cited. And the fragility travels with the result: a p of 0.0138 that becomes 0.057 when one protein leaves is reported as exactly that, not as "p < 0.05".

This is the same standard Rosetta Q asks of everyone else in our checklist for judging an advantage claim and in what radical reproducibility means: a result that only survives the control you happened to pick is not a result yet. A ledger that only publishes the numbers that flatter it is not a ledger — the same reason an advantaged solve is defined so that it can be reported as zero.

What we know and what we don't know

What we know, as of 2026-09-04. The predicted pockets and the real allosteric sites occupy different distance shells (percentile 11 vs 28, Mann-Whitney p < 1e-5). Under a null matched to ±2 Å of the real site's distance, the aggregate result still crosses 0.05 (0.0138 and 0.0107 across two independent seeds, 0.0011 and 0.0076 at the ±1 and ±3 Å bands) but falls about three orders of magnitude from the unmatched 5.4e-6. Removing one protein takes it to 0.057. Individually, 5 of 86 significant against 4.3 expected by chance is not an excess (binomial p = 0.43).

What we don't know. Whether anything survives once the periphery bias is corrected in the detector itself, which is now the top of our queue: as things stand, the primary metric and the discovery queue are partly measuring the same thing, proximity to the active site. Whether the aggregate crossing is signal or two lucky proteins — that needs a bootstrap over proteins, which we have not run. How much of the remaining effect is real, given that the matching is imperfect in a direction that flatters us: we match the seed residue, while the null pocket is the seed plus its k-1 neighbours, so the null pockets end slightly farther out than the real site (EGFR: real 8.3 Å, historical null 16.5 Å, matched null 10.3 Å). The matched p is therefore still somewhat optimistic, not conservative. The matching is on the median distance, not on the full distribution. The effective resolution is set by the pool size (21 to 164 distinct seeds), so the 0.0005 values read as "below the resolution of the test", not as small p-values. And the matched null conditions on the answer — it uses the real site's distance to define the band — which is legitimate for this question but is not blind prediction and must never be presented as such.

One more thing we do not have, and it is ours to fix: the raw artifacts of this run are internal. By the standard Rosetta Q publishes under, a run report should ship with a notebook someone else can execute, and this one ships as a summary. That is a debt, not a feature, and it is written here so it stays visible.

Rosetta Q publishes verdicts with reproducible raw data. This is educational content, not a product claim.

Sources:
· Kapoor & Narayanan — Leakage and the reproducibility crisis in machine-learning-based science, Patterns 4, 100804 (2023)
· Kapoor & Narayanan — preprint version (arXiv:2207.07048)
· Whalen, Schreiber, Noble & Pollard — Navigating the pitfalls of applying machine learning in genomics, Nat. Rev. Genet. 23, 169 (2022)
· Chen et al. — Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening, PLOS ONE 14(8): e0220113 (2019)
· Mysinger, Carchia, Irwin & Shoichet — Directory of Useful Decoys, Enhanced (DUD-E), J. Med. Chem. 55, 6582 (2012)
· Rosetta Q engine run report RQ-PRIMARY-MATCHED, sealed 2026-09-02 — internal; raw artifacts not yet published
· Rosetta Q — Why does a weak classical baseline ruin a quantum benchmark?
· Rosetta Q — What does radical reproducibility mean in benchmarking?
· Rosetta Q — How to tell if a quantum advantage claim is credible
· Rosetta Q — What is an advantaged solve and why does it matter?

Get the next one.

One email a week: what the evidence engine published and what moved in the registry. No pitches.