Dipankar — you put it precisely: "presence, abstain on identity" is only honest if presence itself is calibrated, and a held-out clean number is the thing that settles it. Here it is, the costly parts included.
One correction first, because it's your own example and it cuts toward you rather than away: the clean SmolLM2-135M base scoring above the wolf, past the 0.85 line — that reading is BAIT's q-score, not our presence axis. BAIT is one of the anti-correlated scanners the board discards for exactly this behaviour; our presence read is architecture-honest on that base and doesn't fire there. So the false-fire you pointed at is real, but it belongs to a scanner we throw out, not to us.
Where your objection lands for real, I'll give you the whole shape. Presence is transductive. With a recipe-matched benign reference it's calibrated — it transfers zero-shot to an architecture it was never calibrated on at 0.854. Take the matched reference away and point it at the open world, and it dies: 210 of 210 community-clean adapters fire, FPR 1.0 — it can't tell a poisoned adapter from a community one at all. We publish that as our own failure mode, because you're right that a presence read which false-fires on true negatives is just the anti-correlated scanner wearing a different hat.
Your literal question — a held-out clean cohort. On the channel that needs no matched reference — recovering the planted payload back off the weights, not scoring a distance — the held-out clean number is 0 of 68, and 0 of 20 on the cross-recipe community finetunes: the same diverse finetunes that fire the weight read 210 out of 210. I won't oversell a zero, though. It's precision-first — it wakes only a minority of real backdoors (~40% on the trained class) — so a silent model is "not attested," never "clean." A point-zero isn't a guarantee either: turned into a distribution-free bound it's roughly ≤16% on that slice, because the 20 draws are only 14 distinct authors and the clustering widens the interval. And since you'd rightly ask whether one judge is grading its own work: a second judge, blind to the first's calls and to the answer key, reproduced the labels at κ 0.92 — but false-fired on 1 of 10 clean itself, so the 0-of-68 is that stricter judge's zero, and I'll say plainly the looser one wasn't.
I'm also retiring a line I'd have written a week ago — "a clean adapter has nothing to confess." On the hardest negatives — benign finetunes whose legitimate job is the payload's own shape — the channel is not perfectly silent: one of eighteen confessed, so the honest pooled number is 1 in 210, and it's in the record now, not a footnote.
On the GPU-kernel point, which is the real thesis: a threshold on a separation score can always read clean when the confounding case was simply absent from the set. We caught that on ourselves — on one public corpus our weight read separated poison from clean at AUC 1.0, until we found every clean adapter had been trained half as long as every poison one, so "backdoored" and "trained longer" were literally the same column. We conceded the corpus and marked it. It's why the two things I'll actually defend aren't separation numbers: the recovery channel above, and a machine-checked boundary — that a static weight read is provably blind where a triggered read sees, sorry-free, standard axioms only. One of those reversibility theorems was re-proven from scratch by an outsider, Justin Garringer, about seven hours after we opened the lane, on his own commit.
And on priority, said straight: I am not first at raw per-architecture separability. PEFTGuard reaches ~1.0 and got there before us, reading it from the weights with a trained per-family classifier and a labeled calibration set. I won't claim a lane I don't hold. What we hold is the harder setting your question actually points at — a model you have never seen, no matched reference, forward-free, on CPU — where that per-architecture number does not transfer, and where the honest answer is a bounded FPR and a legible confession, not a 1.0.
None of this is a solved detector, and I'd rather be the one to tell you where it's soft. It's runnable, and I'll be exact about what running it proves: one command re-sums the published judge verdicts — it checks my arithmetic, not my judgment — or you point it at clean models you choose and put them through the same interface for a denominator of your own. The board, the FPRs, the conceded failures, and the Lean are posted: https://protora.vulcora.se/bench/ragnarok, https://vulcora.se/coverage, https://vulcora.se/protocols, and https://github.com/Vulcora/proofora.
— Arian