2026-10-10 · test-retest across sessions
Judge drift report (synthetic demo)
Session A (verdicts-2026-09-28.jsonl, dated 2026-09-28) against session B
(verdicts-2026-10-05.jsonl, dated 2026-10-05), 7 days apart, run 1
of each file.
Summary
Across sessions frontier gave the same class to 52 of 60 paired items (0.867, Wilson 95% [0.758, 0.931]). Its two runs within session A agreed on 55 of 60 (0.917, [0.819, 0.964]); the across-session rate sits inside that interval.
Synthetic demo data from a fixed pattern; not a measurement of any judge. demo/data/, written by demo/data/make_synthetic.py
Drift
Same judge, same items, two sessions
| runs | kind | agree | paired / unpaired | rate [95%] | κ [bootstrap 95%] |
| frontier A vs frontier B | A vs B | 52/60 | 60 / 0 | 0.867 [0.758, 0.931] | 0.787 [0.634, 0.918] |
|---|
| frontier A#1 vs frontier A#2 | within | 55/60 | 60 / 0 | 0.917 [0.819, 0.964] | 0.866 [0.744, 0.972] |
|---|
This is test-retest across sessions, the design a registered judge-reliability study (October
2026) used for drift. Anything that changed between the two days lands in the across-session
figure, sampling noise included. Where a file holds two runs, the within-session row shows how
much of that the judge produces on a single day.
Confusion
frontier A (rows) by frontier B (columns), 52 of 60 on the diagonal (underlined); paired 60 / unpaired 0 | PASS | PARTIAL | FAIL |
|---|
| PASS | 26 | 2 | 2 |
|---|
| PARTIAL | 1 | 16 | 1 |
|---|
| FAIL | 2 | 0 | 10 |
|---|