I could predict my judge's verdict without showing it the answer
2026-07-27 · solo
We use a model to judge whether a diagnostic scenario is any good — realistic, clear, well-formed. It is fast, it is consistent, and it feels like measurement.
Then we ran a test on it, and I want to write down what happened, because it took me by surprise and the test costs almost nothing to repeat.
We trained a small model to guess what the judge would say. Not by reading the scenario the judge was reading — by looking only at *what kind of task had been ordered*. No answer, no text, none of the thing being judged.
It guessed well. Given two scenarios, one the judge passed and one it failed, it picked the right one far more often than chance — while never seeing the thing being judged at all.
I am deliberately not putting a percentage on that here. We have an internal rule about which numbers are allowed to leave the building: a figure only goes public if it is anchored in execution, frozen before it was measured, independent of the gate it is describing, and reproducible by a stranger. This one is a good internal signal and it fails that bar — it points, it does not prove. Publishing it as a headline would be exactly the move this post is about.
What matters isn't our number anyway. It is that the test is cheap, and that you can run it on your own judge this afternoon.
What that meant for us
If our judge's verdict could be guessed from the order slip, then part of what it was "measuring" had been decided before anything was written. In our data, some kinds of task were simply likely to pass, and others were nearly doomed. Our judge was partly reading the work and partly reading the category.
I don't know whether that is true of judges in general, and I'm not going to claim it is. I know it was true of ours, in our data, on the day we looked.
What it changed was how we read our own number. We had been treating it as a measurement of the work in front of it. It wasn't. It was a measurement of the work plus where it came from — and from the outside, you cannot see the split.
The part that fooled me
Our judge score was reproducible. Ask the same model the same question at the same settings and the same number comes back. It felt rigorous. Rerunning it felt like verifying it.
It wasn't. Rerunning told me the measurement was stable. It told me nothing about whether it had measured the right thing. Our blind spot reproduced perfectly, and that consistency read as confidence.
We learned this once the expensive way, on ourselves. We watched a training number climb, clearly and repeatedly, and nearly called it a win. It wasn't. The thing we actually cared about — does the model produce a fix that works — had not moved. We had improved the score, not the product.
So we ask a different question
We don't put a judge score on the front page. Not because judges are useless — ours does a real job filtering for realism and clarity — but because it cannot carry the weight of a claim about whether something works.
The question we use instead is one an answer cannot argue with:
Does the broken version fail the hidden verifier, and does the fixed version pass it, when you run them?
Not does this look correct — does this behave correctly. There is no partial credit, no opinion in the loop, and nothing for a model to be consistently wrong about. A scenario only counts if both sides execute the way they must.
The verifier and the harness go out with the data, so the number can be re-run by whoever is looking at it. That part isn't generosity — a number nobody else can run isn't a number, it's a claim.
The test, if you want to run it on yours
I'm not going to tell you what your judge is doing. I don't know. But the test is cheap, and you can point it at your own:
Hide the thing being judged, and see how well something else predicts the verdict anyway. Metadata, category, source, length — whatever you have that isn't the work itself. Then look at how well it does.
If it does poorly, that is worth knowing and you have lost an afternoon. If it does well, you have learned something about your own number that nobody was going to tell you.
That is the whole reason I'm writing this down. Not to say anything about judges in general — I only looked at one. But it hadn't occurred to me to run this until we did, and if it hasn't occurred to you either, now it has.
We didn't change our policy afterwards; we already refused to headline a judge score. What changed is how quickly we say it out loud when someone asks why our front page doesn't show one.