A frontier-tier model judge read 197 case files in two blind sessions hours apart and gave the same answer on 178 of them: 0.904, with a 95 % interval from 0.854 to 0.937. Half of that disagreement, about ±0.05, is judge noise, and any judged score carries it beyond its sampling interval.
The question
Most judged benchmarks in my programme rest on a model judge, and so does a self-hosted decision service I am testing. Earlier on 6 October a frontier-tier judge classified the right next step in 197 blind checkpoint packets and matched a programmatic oracle on 99 of 120 disputed checkpoints: 0.825, with an interval from 0.747 to 0.883. Nobody had checked whether that judge would say the same thing twice. I also wanted to know whether a cheaper model from the same family reads the packets the way the frontier tier does.
What was registered
Four hypotheses went on file at 23:18 CDT on 6 October, before any verdict existed. The retest would agree with the first session on at least 0.90 of packets, with a kill line at 0.80. The mid tier would reach 0.80 against the frontier tier and the small tier 0.70, with mid above small. At least 0.60 of the frontier tier's self-flips would sit on the boundary between closing a case and interpreting what is on file. No tier would choose BLOCKED on more than 5 of the 197 packets.
How it ran
The 197 packets were copied byte for byte apart from the title line, given new random ids and shuffled. A judge picks one of four classes for the next step: INFO, REASON, DONE or BLOCKED. Three tiers of one vendor's model family judged every packet blind, with eight agents per tier, for 591 verdicts. One small-tier judge wrote a row under an id that belongs to no batch. That row went to a rejected folder and a fresh agent judged the packet it had skipped.
What was found
The retest landed at 0.904, on the registered threshold, with κ 0.827. On the 120 disputed checkpoints behind the 6 October headline, the two frontier sessions agreed on 118, so that result stands without a reliability caveat.
Thirteen of the 19 self-flips sat in the 37 packets where the service had claimed completion and the scorer marked the claim wrong. There the two sessions matched on 24 of 37, or 0.649, with an interval from 0.488 to 0.782.
The mid tier matched the frontier retest on 180 of 197: 0.914, interval 0.866 to 0.945, κ 0.850. Against the first frontier session it reached 0.954, interval 0.915 to 0.976, closer than the frontier tier came to its own earlier run.
The small tier matched the frontier retest on 161 of 197: 0.817, interval 0.757 to 0.865, κ 0.689. It called 30 of 197 cases complete, where the frontier retest did so 23 times and the mid tier 20. None of its 197 verdicts carried a low-confidence flag; the frontier retest flagged 9 and the mid tier 4.
The BLOCKED hole
The 6 October session never chose BLOCKED and read all 33 oracle-BLOCKED cases as INFO or REASON. The fresh frontier session chose it 10 times, six of them on packets the oracle also calls BLOCKED, so part of that hole was a habit of one judging session. The mid tier chose BLOCKED 3 times and matched the oracle on all 3. The small tier chose it once and missed.
Ten of the 19 self-flips went into BLOCKED. That sank the boundary hypothesis, which came in at 9 of 19, or 0.474, and the BLOCKED limit for the frontier tier.
Against the scorer
All three tiers agreed with the programmatic oracle on 0.675 to 0.711 of packets and missed it in the same two places. Of the 53 oracle-DONE cases the frontier tier sent 26 to REASON, the mid tier 33 and the small tier 19. Of the 33 oracle-BLOCKED cases they sent 25, 28 and 19 to INFO. The judges agree with each other more closely than any of them agrees with the scorer, which points first at a rule-table definition that judges and scorer read differently.
What changes
Every judged result I publish now carries two lines beside it: the sampling interval and judge noise of about ±0.05. The mid tier becomes the default judge for the next judged studies, with the frontier tier kept for adjudicating disagreements between judges. The small tier can pre-screen, provided every DONE it issues goes back to a larger judge. Before a study counts BLOCKED, the packet has to state the request budget and the fact that nothing usable is on file.
What it does not claim
Each tier judged once, in one model family. The task was gate classification over two synthetic processes. The two frontier sessions ran hours apart, so prompt order and model sampling are confounded. The oracle is itself disputed on DONE and BLOCKED, so agreement with it measures concordance with the scorer. The ±0.05 belongs to this task, and other tasks may carry more or less.
How it was verified
A second agent recomputed every figure from the raw verdict files with its own code, without reading the first agent's output, and matched every headline to three decimals at 23:32 CDT.
Reproducibility
The registration exists and predates the first verdict. Code, packets and verdict files are available on request until the public repository opens.