Dinand Tinholt · Experiments

Judge Calibrator · library and demo reports

Both reports run on synthetic demo data from a fixed pattern; no model judged anything there.

Python library · standard library only · version 0.2.1

The Judge Calibrator

judgecal runs one or more model judges over a set of items, twice or more, and tells you how often each judge agrees with its own earlier run and with the other judges. It prints the judge-noise band to quote beside any judged benchmark number, and writes a one-page HTML report.

Open a demo report

Read the noise band

The band is half the share of paired items on which the reference judge changed its class between two runs. Rounded to two places, it goes on its own line beside a judged result:

completion rate 0.62 [0.53, 0.70] (Wilson, n = 120); judge noise ±0.05  (illustrative numbers)

The bracket is sampling error from the number of items. The judge's own wobble goes in the band, so it sits apart on the line, after the semicolon. Every agreement figure counts only the items where both sides have a readable verdict, and the report prints paired and unpaired counts beside each one.

Run it on your own judge

Write a rubric file with at least two classes, write the items as JSON lines with an id and a text, and point the tool at any endpoint that speaks the chat-completions protocol, hosted or local:

python3 -m judgecal report --items items.jsonl --rubric rubric.json \
  --endpoint big   https://api.example/v1   big-model-name  JUDGE_API_KEY \
  --endpoint cheap http://localhost:8000/v1 small-model     - \
  --repeats 2 --html report.html --verdicts-out verdicts.jsonl

python3 -m judgecal compare verdicts-2026-10-01.jsonl verdicts-2026-10-08.jsonl \
  --rubric rubric.json --html drift.html

Saved verdicts replay later at no cost. Strata you define get their own agreement figures, a third run shows whether the band moves, and prices you pass turn returned token counts into a cost line.

What it does not do

It reports how consistent judges are with themselves and with each other. Scoring a judge's accuracy takes a ground truth, and the tool has no input for one. Two judges can agree on every item and both be wrong. A flip has many possible causes, and sampling temperature, prompt order and a changed model all land in the same retest number.