Dinand Tinholt · Experiments

Judge Calibrator · drift report

2026-10-10 · test-retest across sessions

Judge drift report (synthetic demo)

Session A (verdicts-2026-09-28.jsonl, dated 2026-09-28) against session B (verdicts-2026-10-05.jsonl, dated 2026-10-05), 7 days apart, run 1 of each file.

Summary

0.867frontier agrees with itself across sessions on 52 of 60 items.Wilson 95% [0.758, 0.931], κ 0.787; paired 60 / unpaired 0

Across sessions frontier gave the same class to 52 of 60 paired items (0.867, Wilson 95% [0.758, 0.931]). Its two runs within session A agreed on 55 of 60 (0.917, [0.819, 0.964]); the across-session rate sits inside that interval.

Synthetic demo data from a fixed pattern; not a measurement of any judge. demo/data/, written by demo/data/make_synthetic.py

Drift

Same judge, same items, two sessions
runskindagreepaired / unpairedrate [95%]κ [bootstrap 95%]
frontier A vs frontier BA vs B52/6060 / 00.867 [0.758, 0.931]0.787 [0.634, 0.918]
frontier A#1 vs frontier A#2within55/6060 / 00.917 [0.819, 0.964]0.866 [0.744, 0.972]

This is test-retest across sessions, the design a registered judge-reliability study (October 2026) used for drift. Anything that changed between the two days lands in the across-session figure, sampling noise included. Where a file holds two runs, the within-session row shows how much of that the judge produces on a single day.

Confusion

frontier A (rows) by frontier B (columns), 52 of 60 on the diagonal (underlined); paired 60 / unpaired 0
PASSPARTIALFAIL
PASS2622
PARTIAL1161
FAIL2010