Mixed2,681 checkpoints across 794 cases in the main retrained-head run
A retrained open decision head scored 0.728 against 0.566 for a rule table on 800 fresh cases, a gap that excludes zero; the first round missed, with 0.545 against 0.575 (gap -0.029, interval +0.001 to -0.060).
Registration and verificationPredictions on the record before any evaluation case was opened; evaluation sets frozen before either run; protocol reviewed independently three times before the first run and once before the second. Both runs rescored from the raw case files by a party who did not build the scoring code. The wrong-claim fix was found after the registered bar was missed, and the piece reports both readings.
The question
Every governed process has a moment where a case could go several ways and something has to pick one: whether we have what we need, whether to ask for something, whether someone should think harder about what is already here, or whether to close it. TypeSafe, a vendor, built a hosted product called Jev around this kind of decision; it answers typed yes/no or multiple-choice questions with a calibrated probability. This campaign asked what happens when a company builds the same idea on its own hardware, and whether an open decision head can finish more cases correctly than a hand-written rule table and than the hosted service.
What was registered
The test compared five ways of choosing the next step in a synthetic supplier-onboarding workflow: a rule table, a local model choosing freely, the same local model asked the identical typed questions, the open decision head, and Jev reached through OpenRouter. All three local approaches ran on the same model, DeepSeek V4.1 Flash, so any difference between them comes down to how the question was asked. The test measured whether each approach can tell "we need a document" apart from "we need to think harder", and whether acting on that distinction finishes more cases correctly.
The predictions were on the record before anyone opened an evaluation case. The evaluation set, eight hundred held-out cases plus a two-hundred-case harder set with shifted formats and a policy change, was built and frozen before either run. An independent review checked the protocol three times before the first run was allowed and once more before the retrained head's run. The registration also set a level for the rate of wrongly calling a case complete, as the trigger for redesigning a safeguard, and a bar of 0.08 for expected calibration error.
What was measured
In the first round (J1) the zero-shot open head, Jev and the rule table ran on 800 eval cases, the harder 200-case set and a development set. In the second round (J5) a retrained head ran on 800 fresh cases (eval2), which logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case, and on the harder set. A follow-on test (J6) carried the same approaches, Jev included, onto a second synthetic process, accounts-payable exception routing, and measured the fast mode under load on two DGX Sparks running four benchmark streams at once. Repeating the same decision twenty times measured consistency. Afterward a party who did not build the scoring code rescored both runs from the raw case files.
Jev's own figures were measured on 24 September 2026 through OpenRouter: 0.35 seconds for a single question at $0.0000128, for 304 input tokens.
What it found
The first round missed. On 800 eval cases the open head scored 0.545 against the rule table's 0.575, a gap of -0.029 with an interval from +0.001 to -0.060, inconclusive. Jev scored 0.540 on the same set, 0.560 on a development set and 0.536 on the 200-case set built with contradictory evidence, and its own gap against the rule table excludes zero on both sets. The rule table finished ahead of both the zero-shot open head and the hosted service.
The second round reversed that. The retrained head scored 0.728 on the 800 fresh cases against 0.566 for the rule table, and 0.739 against 0.565 on the harder set, which put it 0.199 ahead of Jev's 0.536 on the 199 cases both scored. Every one of those gaps excludes zero. A smaller fine-tuned head scored 0.701 on the fresh set, short of the bar for replacing the retrained head, and ships with an experimental label.
The wrong-claim rate came in at 0.047 for the retrained head on the fresh set, above the level registered in advance. Tracing it showed the environment building the evidence pack was listing extra supporting documents it never checked, mostly decoys planted to catch a careless citation. Once it checks every document the way the rest of the system does, the rates are 0.0088 for the rules, 0.0113 for the retrained head and 0.0163 for the learned head. Every case that changed went from wrong to verified and no other case moved. The registered bar is met on the fixed reading and missed on the original one; the fix was found after the bar had been missed. On the contradictory-evidence set in the first round, every approach including the rule table wrongly called a case complete 21 to 28 percent of the time under the original scoring rule, falling to between 1.5 and 9.5 percent under a corrected rule, because the test sometimes offers two equally valid documents and marked the second as wrong.
Speed went to the hosted service. Under four benchmark streams the open head's fast mode had a median response time of 3.1 seconds and a p90 of 4.8, against Jev's 0.35. A gate fires once per checkpoint, and the work around it takes minutes to days, so the gap changes little in most workflows.
Calibration was poor for the first head: expected calibration error near 0.52, roughly six times the registered 0.08 bar. A later head was well calibrated: item-level error 0.0178 uncalibrated and 0.0143 refit on the full eval2, and 0.0313 uncalibrated on the 400 shared cases, against 0.0057 uncalibrated and 0.0054 refit for the Laya encoder on the same 400.
Consistency depended on the serving stack. In the first run, with a second job sharing the endpoint, Jev changed its answer on 0.05 percent of repeats and the open head on 4.95 percent (eval) and 5.85 percent (challenge). Re-run alone, the open head changed on 0.05 percent at one call in flight on both processes, matching Jev, and on 3.4 percent (onboarding) and 1.95 percent (accounts payable) at four calls in flight. The shipped service now defaults to one call in flight.
On the second process the head did not carry over. It scored 0.650 (ap-eval) and 0.659 (ap-challenge) against the rule table's 0.648 and 0.638, level with each other; Jev scored 0.636 and 0.622. A learned head trained on the first process scored 0.522 and 0.558 there untrained, and 0.663 and 0.671 once retrained on 100 of the new process's own labelled cases (271 decisions). Stripped of the judgement head's answers, in training and in serving, it scored 0.660 against 0.663 on one split and 0.670 against 0.671 on the other.
Two open encoders were also tested. Fine-tuned on 1,567 labelled decisions from 363 onboarding cases, Laya scored 0.815 on the harder set (0.824 on a 400-case eval2 subsample) against the rule table's 0.565; on accounts payable, with 201 labelled decisions from 58 cases, it fell to 0.554 (ap-challenge) and 0.539 (ap-eval), below the rule table. CLM-8B scored 0.687 and 0.641 on the onboarding sets and 0.568 and 0.586 on accounts payable. Untuned, both matched the rules and went no further.
Headline numbers, as published
First round: open head decision-class accuracy, 800 eval cases, zero-shot
0.545
First round: rule table decision-class accuracy, 800 eval cases
0.575
First round: open head minus rule table, 800 eval cases (inconclusive)
-0.029[+0.001 to -0.060]
Jev (hosted) decision-class accuracy on the same 800 eval cases
0.540
Jev decision-class accuracy on the 200-case contradictory-evidence set
0.536
Second round: retrained head decision-class accuracy, 800 fresh cases (eval2)
0.728
Second round: rule table decision-class accuracy, 800 fresh cases (eval2)
0.566
Second round: retrained head decision-class accuracy, harder set
0.739
Second round: rule table decision-class accuracy, harder set
0.565
Retrained head ahead of Jev's 0.536 on the 199 cases both scored, harder set (gap excludes zero)
0.199
Smaller fine-tuned head decision-class accuracy, fresh set (experimental label)
0.701
Wrong-claim rate, retrained head, fresh set, before the evidence-pack fix (above the registered trigger)
0.047
Wrong-claim rate, retrained head, after the evidence-pack fix
0.0113
Wrong-claim rate, rule table, after the evidence-pack fix
0.0088
Wrong-claim rate, learned head, after the evidence-pack fix
0.0163
Expected calibration error of the first head (registered bar 0.08)
about 0.52
Item-level calibration error of the later judging head, full eval2: uncalibrated 0.0178, refit
0.0143
Jev single-question response time at $0.0000128 for 304 input tokens, measured 24 September 2026
0.35 seconds
Open head fast-mode median response time, four benchmark streams
3.1 seconds
Open head fast-mode p90 response time, four benchmark streams
4.8 seconds
Checkpoints logged in the J5 eval2 run, across 794 cases (about 3.4 per case)
2,681
Share of repeats that disagree, Jev, first run (open head 4.95% eval, 5.85% challenge, with a second job on the endpoint)
0.05%
Share of repeats that disagree, open head alone at one call in flight (both processes)
0.05%
Share of repeats that disagree, open head at four calls in flight, onboarding (accounts payable 1.95%)
3.4%
Second process (accounts payable): open head accuracy on ap-eval, against rule table 0.648
0.650
Second process: learned head trained on the first process, no retraining, ap-eval
0.522
Second process: learned head retrained on 100 of the new process's labelled cases (271 decisions), ap-eval
0.663
Laya fine-tuned on 1,567 labelled decisions, onboarding harder set (rule table 0.565)
0.815
Laya fine-tuned on 201 labelled decisions from 58 accounts-payable cases, ap-challenge, below the rule table
0.554
AUC of a vote across several readings as a wrong-claim guard (not in use)
0.56
AUC of a guard that read the cited document itself (not in use)
0.52
What it does not claim
The result is not that a company's own hardware beats a well-funded vendor's model at this task in general. The first version of the open head did not beat the rule table, and the first-round gap sits inside its interval.
Three guards for the wrong-claim failure were tried and none is in use. A vote across several readings separated a wrong claim from a right one at an AUC of 0.56, and a version that read the cited document itself reached 0.52, no better than a plain deterministic check. The corrected rate comes from the evidence-pack fix, which was found after the bar was missed.
Two synthetic processes built the same way, with the same mix of case types. Whether any of this holds on a process built differently is open.
The head's lead from the first process was gone on the second. A learned head needs roughly a hundred labelled cases of the new process.
Consistency at four calls in flight was 3.4 and 1.95 percent, and where all twenty repeats agree, more than half of that agreement lands on the wrong answer, for both approaches.
An abstention band abstained on 1.8 and 0.7 percent of claims on the two accounts-payable splits, too few to separate its effect from rerun noise, so that test is inconclusive. A replay test meant to rank Jev's actions against the rule table's favors the rule table by construction and will not be reported as a ranking.
The deliberate eight-sample mode, roughly eight times the cost of a fast call, has not been timed. The J6 run banked no worst-case latency or request count. A phase using a stronger outside model, under a sixty-dollar ceiling, has not run.
The code is public at https://github.com/dtinholt/decision-head under Apache-2.0, after two independent security reviews (23 September and 6 October 2026).