Pillar D · State as of 2026-09-07

What does a frozen protein detector score on structures it has never seen? Week 37: 0 of 1.

Week 2026-W37: 1 eligible structure, 0 blind hits. On 2026-09-07 Rosetta Q froze its de-novo allosteric-site detector with a recorded hash (08ad337d, 86 training proteins, 10 annotation-free features) and began a weekly public scoreboard. Of 165 PDB entries released in the previous seven days, one passed every filter — human phenylalanine hydroxylase with belinostat, entry 9T1O — and the frozen model missed it by 7.2 angstroms while all five of its internal confidence checks passed. The cumulative prospective scoreboard opens at 0 of 0 because the freeze is today; the case is logged on a separate blind-retrospective line at 0 of 1, 95% Wilson interval 0 to 79%.
→ Leer en español
State as of: 2026-09-07
PILLAR D · RUN REPORTRUN DATE: 2026-09-07 · WEEK 2026-W37What does a frozen protein detector scoreon structures it has never seen?0 / 1blind hits by the frozen detector, week 37cumulative prospective scoreboard opens at 0 / 0WEEK 2026-W37 · THE FUNNEL01165 PDB entries released in 7 days0222 passed the ligand filter0310 unique proteins041 eligible: PAH + belinostatblind hits: 0FROZEN MODEL HASH 08ad337d · TRAINED ON 86 PROTEINS · NOT RETRAINED

What did Rosetta Q measure this week?

That its own allosteric-site detector, frozen this morning with a recorded hash, scored 0 of 1 on the only structure released in the past seven days that it was eligible to be tested on. The cumulative prospective scoreboard opens at 0 of 0, because the freeze itself happened today.

This is the first run of a weekly routine with one purpose: to measure how often a detector that cannot be adjusted lands on the right site in proteins that were not in its training set. A model evaluated on the data that built it always looks better than it is. The only number an outside lab can take seriously is the hit rate on structures published after the model was frozen. If that number is low, it gets published low. That is the product.

ONE WEEK OF THE PDB, FILTERED (LOG SCALE)165released · X-ray/cryo-EM · ≤3.0 Å22small-molecule ligand · model organism10unique proteins after dedup1ligand distal to the active site0blind hitssource: RCSB search API · window 2026-08-31 to 2026-09-07queried 2026-09-07

What exactly was frozen?

Not the code — a set of numbers. The examinee is ten features (degree, closeness to the active site, coordination, ClinVar variant density, evolutionary conservation, GNM and ANM mobility, classical diffusion, GNM coupling, betweenness), a scaler, a logistic regression's coefficients and intercept, and the protocol constants: pocket size k = 8, contact cutoff 8.5 Å, distality threshold 6 Å, 2000 permutations, seed 20260717. Those numbers plus the SHA-256 of each code file that produces them hash to 08ad337db711d7e7. Recomputing that hash is how a future run proves the model was not quietly adjusted.

It is deliberately the uncorrected model. Rosetta Q has a measured, published bias in this detector: it fires at the 7th to 13th percentile of distance from the active site, while real allosteric sites live around the 28th. Correcting that would be a new model, and a new model is a challenger, not the incumbent. The incumbent gets frozen with its flaws and made to sit the exam.

Component Value Why it is fixed
Features 10, no UniProt annotations annotation features raise the score by recalling UniProt, not by physics; this is the de-novo metric
Training set 86 proteins, list stored in the record fixed since 2026-08-14; 2 of 88 catalog entries never built and are named in the file
Pocket definition top residue + 7 nearest distal neighbours inherited from the validated pipeline, unchanged
Frozen hash 08ad337db711d7e7… recomputed every run; a mismatch halts the science

How was the single case scored?

In an order that can be checked. The apo structure is built first — only the C-alpha contact network, the ligand plays no part. The frozen weights score every distal residue and the pocket is written down with a timestamp. Only then is the ligand read out of the deposited structure and compared. The prediction carries 2026-09-07T13:03:52Z; the reveal carries 13:03:53Z.

The case was PAH, human phenylalanine hydroxylase, entry 9T1O, released 2026-09-02 at 1.94 Å, in complex with belinostat — an approved HDAC inhibitor being examined against a different target. The detector predicted a pocket at A347–A354 plus A194. The drug actually sits on A248–A263 and A322–A377. Nothing overlaps; the closest predicted residue is 7.2 Å from the nearest true contact.

FIVE CHECKS PASSED. THE ANSWER IS STILL WRONG.rarity vs contiguous null0.007PASSrarity vs distance-matched null0.000PASSanchor dilation R = 6 Årank 1PASSanchor dilation R = 8 Årank 1PASSpocket check (geometric proxy)buried 21PASSRESULT: MISS0 of 8 residues touch the real siteclosest predicted residue: 7.2 Å from the nearest true contactPAH (P00439), apo built from 9T1Oprediction stamped before the ligand was read

What is worth knowing about this particular miss?

That every internal confidence check said the prediction was good. Against the historical contiguous null, the predicted pocket scored in the top 0.7% of random pockets. Against the harder distance-matched null it scored above all 2000 of them. Dilating the anchor to 6 Å and to 8 Å — the test that has closed several past candidates — left the pocket at rank 1 both times. The geometric pocket proxy passed too: buried at 21 neighbours against a protein median of 18, and 62.5% hydrophobic.

Five checks, five passes, and the pocket is in the wrong place. The plain reading is that this detector's internal confidence does not predict whether it is right. That is worth one case of evidence and no more, but it is the direction a second and a third case would either confirm or overturn.

The failure also looks exactly like the bias already on file. The predicted pocket is a contiguous stretch of backbone — A347 through A354 — sitting at the 13th percentile of distance from the anchor. Segment, not cavity; periphery, not pocket. That is the same shape of error measured on 15 earlier candidates and on two fresh targets in the 2026-09-02 batch.

Field Value Source
Entry / released / resolution 9T1O · 2026-09-02 · 1.94 Å · X-ray RCSB entry record, queried 2026-09-07
Protein / ligand PAH (UniProt P00439) · belinostat (5OG) RCSB polymer entity and chemical component records
Anchor (active site) A285, A290, A330 — 3 residues UniProt functional annotations mapped by sequence
Prediction (blind, 13:03:52Z) A347, A348, A350–A354, A194 frozen model 08ad337d, this run
Truth (revealed 13:03:53Z) A248–A255, A263, A281, A285–A290, A322–A330, A377 ligand contacts < 4.5 Å in 9T1O
Verdict miss · overlap 0 of 8 · nearest 7.2 Å this run
Sequence identity vs training set 29.6% (closest: FLT3) pairwise alignment against the 86 training accessions

Is the case itself a clean test?

Only partly, and that has to travel with the number. UniProt annotates just three functional residues for PAH — the iron centre His285, His290, Glu330. Belinostat wraps around that centre: six of its sixteen contacts fall within 6 Å of the anchor and ten fall outside, which is what let the case through the distality filter. Its median distance to the anchor is 9.6 Å, the 7th percentile within its own protein.

So this is a ligand adjacent to the catalytic pocket, admitted as distal because the anchor is minimal. It is a weak case, not a clean allosteric one. It counts — the record is append-only and a case that has been sat is not un-sat — but from week 38 any case whose anchor has fewer than five annotated residues is flagged and reported on its own line. Three other proteins were rejected this week under a related rule: an ADP in a DEAH-box helicase and a reaction product in cGAS are catalytic-site occupants, and the pipeline's own catalogue uses exactly those ligands as markers of the active site elsewhere.

THE SCOREBOARDS, WITH HONEST INTERVALSPROSPECTIVE · released after the freeze0 cases. Opens today. First harvest: 2026-09-14.BLIND RETROSPECTIVE · released before the freeze0 / 1 · 95% Wilson 0–79% · identity vs training set 29.6%0%100%Wilson interval, not the normal approximation.Retrospective LOPO reference 32.6% (n=86) is NOT mixed in.

Why does the prospective scoreboard read 0 of 0 when a case was scored?

Because the freeze happened this morning and the structure was released on 2026-09-02, five days earlier. The training set has been fixed since 2026-08-14, so the model demonstrably never trained on it — but there is no verifiable way to prove nobody looked at it in the days between its release and the freeze. A reviewer would be right to ask, and there is no evidence to offer. So the case is recorded on a separate, clearly labelled blind-retrospective line at 0 of 1, and the prospective scoreboard stays empty until next Monday, when the first structures released after the hash existed become available.

The interval on 0 of 1 is a 95% Wilson interval from 0% to 79%. One case tells you almost nothing; the interval is printed so nobody mistakes a single miss for a measurement. Comparison to the retrospective figure — 32.6% blind top-1 hit rate over 86 proteins, leave-one-protein-out — is deliberately kept side by side and never averaged in.

Did the challenger beat the incumbent?

No, and it could not have been shown to this week either. Alongside the frozen model, Rosetta Q runs a challenger that is allowed to improve. This week's challenger, version 0, attacks the known bias without retraining: it takes the frozen model's scores and standardises them within deciles of distance from the anchor, so a residue no longer wins for being close, only for standing out among residues that are equally far.

It also missed, landing at A352 — one residue from the incumbent's pick. It did move outward, from the 13th to the 17th percentile of distance, which is the right direction and nowhere near enough. Score: challenger 0 of 1, incumbent 0 of 1, zero discordant pairs, McNemar undefined. Promotion needs a paired McNemar at p < 0.05, and with fewer than about ten discordant pairs that is arithmetically unreachable, so the test is computed and logged each week and simply not treated as a promotion signal until the accumulated cases can carry it.

Incumbent (frozen) Challenger v0
Top residue A350 A352
Result miss miss
Distance percentile of the pocket 13th 17th
Week 37 score 0 / 1 0 / 1
McNemar undefined (0 discordant pairs)

What we know and what we do not

What we know. The examinee is frozen with hash 08ad337db711d7e7 over 86 training proteins and 10 annotation-free features. In the window 2026-08-31 to 2026-09-07 the PDB released 165 entries meeting the structural filters; 22 carried a small-molecule ligand on a model-organism protein; 10 proteins remained after deduplication; 1 passed the geometric distality check. That case, PAH with belinostat, was predicted blind and missed by 7.2 Å, with all five internal checks passing. Blind-retrospective scoreboard: 0 of 1, 95% Wilson 0–79%. Prospective scoreboard: 0 of 0, opening today.

What we do not know. Whether the detector's confidence measures are uninformative in general or merely on this case — one example cannot separate those. Whether the miss reflects the model or the case, since the three-residue anchor made PAH a weak test to begin with. What the prospective rate will settle at: the measured periphery bias means it will probably start low, and the honest expectation is that it stays below the 32.6% retrospective figure. Whether a properly trained challenger — retrained on an expanded set plus accumulated failures — does better than the cheap residualisation run here; that was not run this week and is the first item of the next one, along with the orthologue and ASD data preload that also did not fit in the window.

Three gaps in the method itself are ours to declare. The pocket-quality check ran as an explicit geometric substitute because fpocket is not installed in this environment, and a substitute is not fpocket. The run payload is written but unsealed: the sealing library is not in the project, and sealing inline is forbidden by house rule, so the artefact sits as a pending-seal file until the library is supplied. And the raw artefacts remain internal, which is the same reproducibility debt this series has declared before and has not yet paid.

Limit statement. A negative from this detector does not authorise anyone to rule out a site. Its blind top-1 hit rate is 32.6% on the retrospective set; a tool that finds roughly one site in three has no power to prove a site is not there. Nothing in this report says belinostat's pocket, or any other, is uninteresting — only that our frozen model did not find it.

Rosetta Q publishes verdicts with reproducible raw data. This is educational content, not a product claim.

Sources:
· RCSB PDB — entry 9T1O, phenylalanine hydroxylase with belinostat, released 2026-09-02 (queried 2026-09-07)
· RCSB PDB — search API, release-date window 2026-08-31 to 2026-09-07 (queried 2026-09-07)
· RCSB PDB — chemical component 5OG (belinostat)
· UniProt P00439 — phenylalanine-4-hydroxylase, functional site annotations
· Wilson, E. B. — Probable inference, the law of succession, and statistical inference, J. Am. Stat. Assoc. 22, 209 (1927)
· Kapoor and Narayanan — Leakage and the reproducibility crisis in machine-learning-based science, Patterns 4, 100804 (2023)
· Rosetta Q run record RQ-PROSP-2026-W37, written 2026-09-07 — internal, pending seal; raw artefacts not yet published
· Rosetta Q — Why can a detector look good just by being close to the active site?
· Rosetta Q — What is an advantaged solve and why does it matter?
· Rosetta Q — What does radical reproducibility mean in benchmarking?

Get the next one.

One email a week: what the evidence engine published and what moved in the registry. No pitches.