
The Judge Calibrator runs one or more model judges over a set of items, twice or more, and tells you how often each judge agrees with its own earlier run. It hands you the judge-noise band, half the share of verdicts that flipped between runs, to print beside any judged benchmark number. It is a small Python library with no dependencies, and the page here opens two reports it wrote from synthetic data.
What it does
Give it a rubric file and a JSON-lines file of items, and point it at any endpoint that speaks the chat-completions protocol, hosted or local. It sends four requests at a time per judge, or it replays verdicts you already saved. Cheaper judges can ride along, and the report shows how often each one matches the judge you trust.
The report opens with a one-page summary that prints on its own sheet. Below it sit Cohen's κ with a seeded bootstrap interval and the confusion tables. A paragraph names the class your judge wavers on most. Strata you define get their own agreement figures. A third run shows whether the band moves, and the token counts the endpoint returns become a cost line.
Save the verdicts, run it again next week, and the compare command reports drift between the two sessions, set beside the judge's own same-day retest.
How to use it
Open the demo reports first. The consistency report and the drift report both run on synthetic data from a fixed pattern, and the three judges are named frontier, mid and small only to show how tiers read on the page.
To run it on your own judge, write a rubric with at least two classes and an items file with an id and a text on each line, then run the report command with one endpoint per judge and two repeats. Quote the band on its own line after the sampling interval, like this: completion rate 0.62 [0.53, 0.70] (Wilson, n = 120); judge noise ±0.05. Those numbers are illustrative. The library's public source link goes on this page when the code is released.
What it does not do
It measures how consistent judges are with themselves and with each other. Scoring a judge's accuracy needs a ground truth, and the tool has no input for one. Two judges can agree on every item and both be wrong.
A flip has many possible causes. Sampling temperature and a model that changed between runs land in the same retest number. Items are pooled, so if they fall into kinds, run each kind as its own file. The endpoint judge has been tested against fake transports and loopback servers, and never yet against a real model endpoint.
Where the numbers come from
The method follows The Judge's Mirror, the registered judge-reliability study the lab ran in October 2026: raw agreement with Wilson 95% intervals and Cohen's κ, with the band set at half the reference judge's test-retest disagreement. No figure from that study appears in the demo reports; its results are on the study page. Every number in the two demo reports comes from synthetic data written by a seeded script, and both reports carry a stamp saying so. The protocol the study followed is on the methods page, and the Error Independence campaign covers the related question of when a second checker shares the first one's errors.
Open the tool
The tool opens on a short landing page with links to both demo reports. It runs in your browser and sends nothing anywhere.
Open the tool Full screen in a new tab Consistency report Drift report