Synthetic demo data from a fixed pattern; not a measurement of any judge.
2026-10-10 · rubric synthetic-demo
Judge consistency report (synthetic demo)
3 judges (frontier, mid, small) classified 60 items into 3 classes (PASS, PARTIAL, FAIL), each judge up to 3 times.
Summary
| ±0.04 | Judge noise to print beside a number this judge produces.half of 5/60 = 0.083 test-retest disagreement, frontier; unrounded 0.042; paired 60 / unpaired 0 |
| 0.917 | frontier agrees with its own second run on 55 of 60 items.Wilson 95% [0.819, 0.964], Cohen’s κ 0.866, bootstrap 95% [0.747, 0.971]; paired 60 / unpaired 0 |
| 0.883 | mid matches frontier on 53 of 60 items, the closest of the other judges.Wilson 95% [0.778, 0.942], κ 0.812; inside the reference’s own retest interval [0.819, 0.964] |
frontier changed its class on 5 of 60 paired items between run 1 and run 2. PARTIAL took part in 5 of those flips, more than any other class. The commonest move ran between PARTIAL and FAIL: 2 went from PARTIAL on run 1 to FAIL on run 2, 2 the other way. Specific agreement, the share of verdicts naming a class that the other run matched, bottoms out at 0.833 for FAIL (24 verdicts named it).
Rounded to two places, frontier's band moves from ±0.04 with 2 runs to ±0.06 with 3. Taken one pair at a time, the 3 pairs give bands from ±0.042 to ±0.075.
Across 2 strata from your strata file, the reference retest is lowest in short cases: 27 of 30 (0.900, Wilson [0.744, 0.965]).
judge noise ±0.04 (half the frontier test-retest disagreement, 5/60 = 0.083; paired n = 60, unpaired 0)
cost: no prices given (pass --price NAME=IN,OUT in USD per million tokens)
demo/data/, written by demo/data/make_synthetic.pyTest-retest
| runs | agree | paired / unpaired | rate [95%] | κ [bootstrap 95%] | band |
|---|---|---|---|---|---|
| frontier#1 vs frontier#2 | 55/60 | 60 / 0 | 0.917 [0.819, 0.964] | 0.866 [0.747, 0.971] | ±0.042 |
| frontier#1 vs frontier#3 | 54/60 | 60 / 0 | 0.900 [0.799, 0.953] | 0.841 [0.712, 0.948] | ±0.050 |
| frontier#2 vs frontier#3 | 51/60 | 60 / 0 | 0.850 [0.739, 0.919] | 0.763 [0.613, 0.895] | ±0.075 |
| mid | one run only, not measured | ||||
| small | one run only, not measured | ||||
Agreement counts only items with a readable verdict on both runs. Unreadable replies and missing items leave the denominator and appear as unpaired. The band is half the share of paired items where a judge changed its class between runs; rounded to two places, it is the figure to print beside any score this judge produced. κ intervals are percentile bootstraps over items, 10,000 draws, seed 1.
Repeat sensitivity
| runs | pairs | flips | band |
|---|---|---|---|
| frontier, runs 1 to 2 | 1 | 5/60 | ±0.042 |
| frontier, runs 1 to 3 | 3 | 20/180 | ±0.056 |
Rounded to two places, frontier's band moves from ±0.04 with 2 runs to ±0.06 with 3. Taken one pair at a time, the 3 pairs give bands from ±0.042 to ±0.075.
Every pair of runs also has its own row in the test-retest table. Pairs of runs share items, so the pooled figure is a description of this run, not an independent estimate.
Where the judge wavers
frontier
frontier changed its class on 5 of 60 paired items between run 1 and run 2. PARTIAL took part in 5 of those flips, more than any other class. The commonest move ran between PARTIAL and FAIL: 2 went from PARTIAL on run 1 to FAIL on run 2, 2 the other way. Specific agreement, the share of verdicts naming a class that the other run matched, bottoms out at 0.833 for FAIL (24 verdicts named it).
| class | run 1 | run 2 | both | specific | flips |
|---|---|---|---|---|---|
| PASS | 30 | 29 | 29 | 0.983 | 1 |
| PARTIAL | 18 | 19 | 16 | 0.865 | 5 |
| FAIL | 12 | 12 | 10 | 0.833 | 4 |
Specific agreement for a class is twice the items both runs put in it, over the verdicts that named it on either run. Flips counts the paired items where one run named the class and the other did not, so each flip is counted under two classes.
Between judges
| pair | agree | paired / unpaired | rate [95%] | κ [bootstrap 95%] |
|---|---|---|---|---|
| frontier vs mid | 53/60 | 60 / 0 | 0.883 [0.778, 0.942] | 0.812 [0.673, 0.923] |
| frontier vs small | 47/60 | 60 / 0 | 0.783 [0.664, 0.869] | 0.653 [0.477, 0.812] |
| mid vs small | 42/60 | 60 / 0 | 0.700 [0.575, 0.801] | 0.520 [0.331, 0.690] |
Cohen’s κ discounts the agreement two judges would reach by chance from their class mix alone. A cheaper judge whose agreement with the reference sits inside the reference’s own test-retest interval is, on this rubric, hard to tell apart from it.
Strata
| comparison | agree | paired / unpaired | rate [95%] | κ |
|---|---|---|---|---|
| long cases · 30 items | ||||
| frontier#1 vs frontier#2 | 28/30 | 30 / 0 | 0.933 [0.787, 0.982] | 0.890 |
| frontier vs mid | 26/30 | 30 / 0 | 0.867 [0.703, 0.947] | 0.779 |
| frontier vs small | 26/30 | 30 / 0 | 0.867 [0.703, 0.947] | 0.781 |
| mid vs small | 24/30 | 30 / 0 | 0.800 [0.627, 0.905] | 0.672 |
| short cases · 30 items | ||||
| frontier#1 vs frontier#2 | 27/30 | 30 / 0 | 0.900 [0.744, 0.965] | 0.843 |
| frontier vs mid | 27/30 | 30 / 0 | 0.900 [0.744, 0.965] | 0.843 |
| frontier vs small | 21/30 | 30 / 0 | 0.700 [0.521, 0.833] | 0.525 |
| mid vs small | 18/30 | 30 / 0 | 0.600 [0.423, 0.754] | 0.380 |
The strata come from the file you supplied; judgecal never derives them from verdicts or answers. Items the file does not mention sit under “(no stratum)”. Small strata give wide intervals; the κ here has no bootstrap interval.
Confidence
| run | low | unreadable | missing | class mix |
|---|---|---|---|---|
| frontier run 1 | 2/60 0.033 | 0 | 0 | PASS 30, PARTIAL 18, FAIL 12 |
| frontier run 2 | 7/60 0.117 | 0 | 0 | PASS 29, PARTIAL 19, FAIL 12 |
| frontier run 3 | 4/60 0.067 | 0 | 0 | PASS 28, PARTIAL 18, FAIL 14 |
| mid run 1 | 5/60 0.083 | 0 | 0 | PASS 31, PARTIAL 15, FAIL 14 |
| small run 1 | 1/60 0.017 | 0 | 0 | PASS 29, PARTIAL 18, FAIL 13 |
Low means the label “low” or a numeric confidence under 50 of 100. Read this column beside the retest table to see whether a judge flags the items it later flips.
Confusion
| PASS | PARTIAL | FAIL | |
|---|---|---|---|
| PASS | 29 | 1 | 0 |
| PARTIAL | 0 | 16 | 2 |
| FAIL | 0 | 2 | 10 |
| PASS | PARTIAL | FAIL | |
|---|---|---|---|
| PASS | 28 | 1 | 1 |
| PARTIAL | 2 | 14 | 2 |
| FAIL | 1 | 0 | 11 |
| PASS | PARTIAL | FAIL | |
|---|---|---|---|
| PASS | 25 | 3 | 2 |
| PARTIAL | 3 | 13 | 2 |
| FAIL | 1 | 2 | 9 |
| PASS | PARTIAL | FAIL | |
|---|---|---|---|
| PASS | 24 | 4 | 3 |
| PARTIAL | 3 | 10 | 2 |
| FAIL | 2 | 4 | 8 |
Cost
cost: no prices given (pass --price NAME=IN,OUT in USD per million tokens)
Cost comes from the token counts the endpoint returned with each reply, at the prices you passed in US dollars per million tokens. Calls that failed, and retried attempts whose reply was unreadable, are left out; a replayed judge shows the counts recorded from its original run.
Limits
- The band follows a convention from a registered judge-reliability study: half the observed test-retest disagreement. Print it on its own line beside the sampling interval.
- A retest figure cannot separate sampling noise from a model that changed between runs; both land in the same number.