Dinand Tinholt · Experiments

Judge Calibrator · consistency report

Synthetic demo data from a fixed pattern; not a measurement of any judge.

2026-10-10 · rubric synthetic-demo

Judge consistency report (synthetic demo)

3 judges (frontier, mid, small) classified 60 items into 3 classes (PASS, PARTIAL, FAIL), each judge up to 3 times.

Summary

±0.04Judge noise to print beside a number this judge produces.half of 5/60 = 0.083 test-retest disagreement, frontier; unrounded 0.042; paired 60 / unpaired 0
0.917frontier agrees with its own second run on 55 of 60 items.Wilson 95% [0.819, 0.964], Cohen’s κ 0.866, bootstrap 95% [0.747, 0.971]; paired 60 / unpaired 0
0.883mid matches frontier on 53 of 60 items, the closest of the other judges.Wilson 95% [0.778, 0.942], κ 0.812; inside the reference’s own retest interval [0.819, 0.964]

frontier changed its class on 5 of 60 paired items between run 1 and run 2. PARTIAL took part in 5 of those flips, more than any other class. The commonest move ran between PARTIAL and FAIL: 2 went from PARTIAL on run 1 to FAIL on run 2, 2 the other way. Specific agreement, the share of verdicts naming a class that the other run matched, bottoms out at 0.833 for FAIL (24 verdicts named it).

Rounded to two places, frontier's band moves from ±0.04 with 2 runs to ±0.06 with 3. Taken one pair at a time, the 3 pairs give bands from ±0.042 to ±0.075.

Across 2 strata from your strata file, the reference retest is lowest in short cases: 27 of 30 (0.900, Wilson [0.744, 0.965]).

judge noise ±0.04 (half the frontier test-retest disagreement, 5/60 = 0.083; paired n = 60, unpaired 0)

cost: no prices given (pass --price NAME=IN,OUT in USD per million tokens)

demo/data/, written by demo/data/make_synthetic.py

Agreement

Agreement with 95% Wilson intervalsfrontier, run 1 vs run 2: 0.917 [0.819, 0.964]; frontier vs mid: 0.883 [0.778, 0.942]; frontier vs small: 0.783 [0.664, 0.869]; mid vs small: 0.700 [0.575, 0.801]0.500.600.700.800.901.00frontier, run 1 vs run 2 · paired 60 / unpaired 00.917 [0.819, 0.964]frontier vs mid · paired 60 / unpaired 00.883 [0.778, 0.942]frontier vs small · paired 60 / unpaired 00.783 [0.664, 0.869]mid vs small · paired 60 / unpaired 00.700 [0.575, 0.801]
Each marker is raw agreement, with its 95% Wilson interval drawn as a bar; the tables below hold the figures. A round ledger-blue marker means a judge was set against its own second run, while the grey squares compare pairs of judges, all on run 1. Where a second judge lands near the dashed line, which marks the reference judge’s test-retest agreement, it agrees with the first about as often as the first agrees with itself.demo/data/, written by demo/data/make_synthetic.py

Test-retest

Each judge against its own repeat
runsagreepaired / unpairedrate [95%]κ [bootstrap 95%]band
frontier#1 vs frontier#255/6060 / 00.917 [0.819, 0.964]0.866 [0.747, 0.971]±0.042
frontier#1 vs frontier#354/6060 / 00.900 [0.799, 0.953]0.841 [0.712, 0.948]±0.050
frontier#2 vs frontier#351/6060 / 00.850 [0.739, 0.919]0.763 [0.613, 0.895]±0.075
midone run only, not measured
smallone run only, not measured

Agreement counts only items with a readable verdict on both runs. Unreadable replies and missing items leave the denominator and appear as unpaired. The band is half the share of paired items where a judge changed its class between runs; rounded to two places, it is the figure to print beside any score this judge produced. κ intervals are percentile bootstraps over items, 10,000 draws, seed 1.

Repeat sensitivity

Band pooled over every pair among the first m runs
runspairsflipsband
frontier, runs 1 to 215/60±0.042
frontier, runs 1 to 3320/180±0.056

Rounded to two places, frontier's band moves from ±0.04 with 2 runs to ±0.06 with 3. Taken one pair at a time, the 3 pairs give bands from ±0.042 to ±0.075.

Every pair of runs also has its own row in the test-retest table. Pairs of runs share items, so the pooled figure is a description of this run, not an independent estimate.

Where the judge wavers

frontier

frontier changed its class on 5 of 60 paired items between run 1 and run 2. PARTIAL took part in 5 of those flips, more than any other class. The commonest move ran between PARTIAL and FAIL: 2 went from PARTIAL on run 1 to FAIL on run 2, 2 the other way. Specific agreement, the share of verdicts naming a class that the other run matched, bottoms out at 0.833 for FAIL (24 verdicts named it).

frontier, run 1 vs run 2, per class
classrun 1run 2bothspecificflips
PASS3029290.9831
PARTIAL1819160.8655
FAIL1212100.8334

Specific agreement for a class is twice the items both runs put in it, over the verdicts that named it on either run. Flips counts the paired items where one run named the class and the other did not, so each flip is counted under two classes.

Between judges

Pairwise, run 1 of each judge
pairagreepaired / unpairedrate [95%]κ [bootstrap 95%]
frontier vs mid53/6060 / 00.883 [0.778, 0.942]0.812 [0.673, 0.923]
frontier vs small47/6060 / 00.783 [0.664, 0.869]0.653 [0.477, 0.812]
mid vs small42/6060 / 00.700 [0.575, 0.801]0.520 [0.331, 0.690]

Cohen’s κ discounts the agreement two judges would reach by chance from their class mix alone. A cheaper judge whose agreement with the reference sits inside the reference’s own test-retest interval is, on this rubric, hard to tell apart from it.

Strata

Agreement within each stratum from the strata file
comparisonagreepaired / unpairedrate [95%]κ
long cases · 30 items
frontier#1 vs frontier#228/3030 / 00.933 [0.787, 0.982]0.890
frontier vs mid26/3030 / 00.867 [0.703, 0.947]0.779
frontier vs small26/3030 / 00.867 [0.703, 0.947]0.781
mid vs small24/3030 / 00.800 [0.627, 0.905]0.672
short cases · 30 items
frontier#1 vs frontier#227/3030 / 00.900 [0.744, 0.965]0.843
frontier vs mid27/3030 / 00.900 [0.744, 0.965]0.843
frontier vs small21/3030 / 00.700 [0.521, 0.833]0.525
mid vs small18/3030 / 00.600 [0.423, 0.754]0.380

The strata come from the file you supplied; judgecal never derives them from verdicts or answers. Items the file does not mention sit under “(no stratum)”. Small strata give wide intervals; the κ here has no bootstrap interval.

Confidence

Low-confidence verdicts and class mix per run
runlowunreadablemissingclass mix
frontier run 12/60 0.03300PASS 30, PARTIAL 18, FAIL 12
frontier run 27/60 0.11700PASS 29, PARTIAL 19, FAIL 12
frontier run 34/60 0.06700PASS 28, PARTIAL 18, FAIL 14
mid run 15/60 0.08300PASS 31, PARTIAL 15, FAIL 14
small run 11/60 0.01700PASS 29, PARTIAL 18, FAIL 13

Low means the label “low” or a numeric confidence under 50 of 100. Read this column beside the retest table to see whether a judge flags the items it later flips.

Confusion

frontier#1 (rows) by frontier#2 (columns), 55 of 60 on the diagonal (underlined); paired 60 / unpaired 0
PASSPARTIALFAIL
PASS2910
PARTIAL0162
FAIL0210
frontier (rows) by mid (columns), 53 of 60 on the diagonal (underlined); paired 60 / unpaired 0
PASSPARTIALFAIL
PASS2811
PARTIAL2142
FAIL1011
frontier (rows) by small (columns), 47 of 60 on the diagonal (underlined); paired 60 / unpaired 0
PASSPARTIALFAIL
PASS2532
PARTIAL3132
FAIL129
mid (rows) by small (columns), 42 of 60 on the diagonal (underlined); paired 60 / unpaired 0
PASSPARTIALFAIL
PASS2443
PARTIAL3102
FAIL248

Cost

cost: no prices given (pass --price NAME=IN,OUT in USD per million tokens)

Cost comes from the token counts the endpoint returned with each reply, at the prices you passed in US dollars per million tokens. Calls that failed, and retried attempts whose reply was unreadable, are left out; a replayed judge shows the counts recorded from its original run.

Limits