The Open Decision Head
Dinand Tinholt, Capgemini, Head of AI Center of Excellence, Americas 23 September 2026

The business question
Every governed process has a moment where a case could go several ways and something has to pick one. In a supplier onboarding queue, an accounts-payable exception or a claims file, the person at that moment is asking whether we have what we need, whether to ask for something, whether someone should think harder about what's already here, or whether to close it. A wrong call there means a case closed early with an open obligation nobody caught, or hours of reasoning spent on a file that was missing one document all along. TypeSafe, a vendor, built a product called Jev around this kind of decision. You hand it a case and a list of yes/no or multiple-choice questions, and it answers each one with a calibrated probability, in 0.35 seconds and for $0.0000128 a question when we measured it. This brief looks at what happens when a company builds the same idea on its own hardware, and at what an operator gets and gives up by doing so.

What the open decision head offers
There are nine advantages, and each one follows from running the service yourself instead of sending a case to someone else's.
Data stays on your network. The case and every answer about it stay on hardware the company already controls, documents included. A pilot can begin before legal has finished a data-processing agreement, since there's no vendor boundary for that agreement to cover.
Calibration you own. The confidence number behind each answer is fitted on the company's own labelled cases, and an auditor can inspect it as a reliability table. When the company's documents change, the calibration gets refit. A hosted service's stated confidence stays as the vendor fitted it.
Customizable in about a day. Adding a use case takes a new question file and a calibration run, and the integration and the contract stay as they are. With a hosted general model, each process goes through its own validation. The open head reaches a second process through the same integration, in about a day.
Fine-tunable, with a measured payoff. Once enough of the company's own decisions have piled up, a smaller model can be distilled on those cases, faster and calibrated for that one process. A first version of this has run, and I report it later in this brief. It scored 0.701 against the full-sized retrained head's 0.728, short of the bar for replacing it, so it carries an experimental label.
Audit per decision. Every answer carries a row with the exact question, the raw probability, the model that answered and how long it took. You can rebuild a decision from six months ago in full.
Base model of choice. The same service runs on the model this project uses today (DeepSeek V4.1 Flash, served as one model spanning two GPUs on the fleet's own hardware). It also runs on a small model on a single workstation, or on a frontier model reached through the company's own cloud account, and the approach stays the same.
Isolation checked on your own data. Jev's premise is that each question gets answered without seeing the others. The open version tests that directly on the company's data and keeps the result where a reviewer can check it.
Cost from the power bill. Running cost comes off the electricity bill with no signup step, and capacity is whatever the company's own hardware provides. The model file is pinned by its own fingerprint, so nobody can quietly swap it underneath a company's own testing.
Open source. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head, after two independent security reviews (23 September and 6 October 2026). A company's security team can read every line that touches its data before the first real case goes through.
The hosted service keeps some advantages. Its model was trained specifically for this kind of typed decision, where the open head reads a general model one token at a time. It publishes a (self-reported) accuracy figure across its own evaluation set, and there's a vendor behind its support line. The experiment below puts a measured number on the gap between the two.
What this looks like in practice (an illustrative smoke test,
outside the benchmark). In the service's own validation run, a
supplier's insurance certificate named the parent company instead of the
subsidiary applying for onboarding. Asked whether the certificate met
the requirement, the service said no, at 99.9 percent probability. Then
it was asked what should happen next, with continue, request more
information and escalate to a person as the options. It picked request
with probability 0.91, correctly reading the gap as one the supplier
could fix. On a three-level urgency scale it answered medium, but only
just, with its probability split almost evenly across all three levels.
That's the kind of case a company would want a person to glance at
before anyone closes it on autopilot. Once the service had looked at the
case once, all three answers came back in under three seconds. The
fleet's governance log for 23 September 2026,
wiki/governance/log-2026-09-23.md, records the satisfied
and request readings.
Why self-hosted matters to a Fortune 500 buyer
Four reasons come up in nearly every procurement conversation about this kind of technology, and each maps to one of the properties above. Data residency and vendor risk review usually take longer than anything else before a pilot starts. A self-hosted service removes the reason for that review to exist. Owning the calibration fit gives a compliance function a confidence number it can defend, where a vendor's number has to be accepted as stated. A vendor waitlist, or a vendor closing enrollment altogether, is a live risk here. TypeSafe stopped issuing new API keys on the day this project started, and that's what triggered building the open version at all. Running on hardware the company already owns turns a per-decision invoice into whatever that hardware already costs to run.
Speed versus accuracy: what matters here
Measured on 24 September 2026, Jev, reached through OpenRouter, answered a single question in 0.35 seconds at a cost of $0.0000128 for 304 input tokens.
The open head's fast mode was tested under real load on the fleet's
own hardware, two DGX Sparks running four benchmark streams at once, and
the results sit in PREREG-J6.md and
results/design-review-J6-2026-09-24.md:
| measure | value |
|---|---|
| median response time | 3.1 seconds |
| p90 response time | 4.8 seconds |
This run didn't bank a worst-case latency or a request count.
Jev is the faster service. A deliberate mode that samples eight answers per decision, meant for harder cases, will cost roughly eight times a single fast call. Nobody has measured its speed yet.
In most enterprise workflow gates the latency gap changes nothing. A gate fires once per checkpoint in a multi-step process. The J5 eval2 run logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case. The work around each gate, such as fetching a document from a supplier or routing a file to an expert, takes minutes to days. A three-second decision inside a step that already waits two days adds almost nothing to the wait, and a wrong call costs far more.

On 800 eval cases, the open head's zero-shot accuracy was 0.545 against a hand-written rule table's 0.575. A statistical check puts that gap of -0.029 somewhere between +0.001 and -0.060. Jev scored 0.540 on the same eval set, 0.560 on a development set and 0.536 on a 200-case set built with contradictory evidence, and its own gap against the rule table excludes zero on both sets. On that first run the rule table finished ahead of both the zero-shot open head and the hosted service, though the open head's gap sat inside the interval. On the contradictory-evidence set, every approach tested, the rule table included, wrongly called a case complete 21 to 28 percent of the time under the original scoring rule. A review afterward traced most of that to a defect in the scoring rule itself, because the test sometimes offers two equally valid documents and marks an answer wrong for citing the second one. Under a corrected scoring rule the rate falls to between 1.5 and 9.5 percent depending on the approach.
A second round retrained the open head on the same task and tested it on 800 fresh cases. It scored 0.728 there against 0.566 for the rule table. The harder set showed the same pattern, with 0.739 against the rule table's 0.565, which put the head 0.199 ahead of Jev's own 0.536 on the 199 cases both scored. Every one of those gaps excludes zero, so this is the first result in the project where the open head beats both the rule table and the hosted service. On the fresh set, the retrained head's rate of wrongly calling a case complete first came in at 0.047. That's above the level the project set in advance as the trigger for redesigning a safeguard. When we traced why, the cause sat in the environment that builds the evidence pack. It was listing extra supporting documents it never checked, most of them decoys planted to catch a careless citation, so the error belonged to the pack builder. Once it checks every document the way the rest of the system already does, the rate drops to these figures:
- rules: 0.0088
- retrained head: 0.0113
- learned head: 0.0163
Every case that changed went from wrong to verified, and no other case moved. The pre-registered bar is met on the fixed reading and missed on the original one. We report both, because the fix was found after the bar had been missed.

For a checkpoint decision I'd rank a gate first on accuracy against the company's own cases and on keeping the data inside the company's network. Next come a model that retrains on a hundred of the company's own labelled cases with no vendor ticket involved and an audit row behind every call, with speed after those. Anyone choosing should weigh Jev's strengths too. It wins outright on latency and on needing no infrastructure, and it learned from a training set far larger than anything one company logs on its own.
What the experiment showed
A registered test, described in a companion technical paper, compared five ways of choosing the next step in a synthetic supplier-onboarding workflow. The first was a rule table and the second a local model choosing freely. Then came the same local model asked the identical typed questions the decision head asks, the open decision head itself, and TypeSafe's own Jev, reached through OpenRouter. All three local approaches ran on the same model, the fleet's own DeepSeek V4.1 Flash, so any difference between them comes down to how the question was asked. The test measured whether each approach can tell "we need a document" apart from "we need to think harder," and whether acting on that distinction finishes more cases correctly.
The predictions were on the record before anyone opened a single evaluation case. In the first run the open decision head landed slightly below the rule table, where we had predicted it would land above the free-choosing model, and Jev landed below the rule table too. The second run tested a retrained head on a fresh set of cases, and that head beat both. The table below reports both runs.

Results.
| what | J1, first run | J5, retrained head, fresh cases | J6, second process (AP) |
|---|---|---|---|
| decision-class accuracy, rule table | 0.575 (eval) | 0.566 (eval2) / 0.565 (challenge) | 0.648 (ap-eval) / 0.638 (ap-challenge) |
| decision-class accuracy, open head | 0.545 (eval), below the rule table | 0.728 (eval2) / 0.739 (challenge), above the rule table | 0.650 (ap-eval) / 0.659 (ap-challenge), level with the rule table |
| decision-class accuracy, Jev (hosted) | 0.540 (eval) / 0.536 (challenge), below the rule table | not re-run on the fresh set | 0.636 (ap-eval) / 0.622 (ap-challenge), at or below the rule table and the head |
| decision-class accuracy, learned head | not applicable | 0.701 (eval2), retrained on the first process, experimental | 0.522 (ap-eval) / 0.558 (ap-challenge) untrained on the new process; 0.663 / 0.671 once retrained on 100 of the new process's own labelled cases |
| consistency, share of repeats that disagree | 0.05% (Jev), 4.95% eval / 5.85% challenge (open head), while a second job shared the endpoint | not re-measured on this arm | re-measured on an isolated endpoint by a follow-on test: 0.05% at one call in flight, matching Jev exactly; rising to 1.95% at four calls in flight |
| wrong-claim rate, retrained head | not applicable | 0.047 (eval2), fixed to 0.0113 once the evidence pack was corrected | not read the same way; see below |
| claim guard | not applicable | not needed once the evidence pack was corrected; two model-based attempts added nothing beyond that fix | not tested on the new process |
| decision-class accuracy, Laya (fine-tuned) | not applicable | 0.815 (challenge) / 0.824 (eval2, 400-case subsample), both above the judging head | 0.554 (ap-challenge) / 0.539 (ap-eval), both below the rule table |
| decision-class accuracy, CLM-8B (fine-tuned) | not applicable | 0.687 (challenge) / 0.641 (eval2 subsample), above the rule table, below the head on eval2 | 0.568 (ap-challenge) / 0.586 (ap-eval), below the rule table |
| item-level calibration (ECE) | not applicable | judging head 0.0178 uncalibrated / 0.0143 refit on full eval2, 0.0313 uncalibrated on the 400 shared cases; Laya 0.0057 uncalibrated / 0.0054 refit on the same 400 cases | not measured on this process |
Every accuracy gap above is a statistically real difference apart from the first run's open head against the rule table, where the interval runs from +0.001 to -0.060 and the result stays inconclusive. J1 also measured whether the head's stated confidence tracked its own correctness, and the first head's calibration was poor there. Its expected calibration error came in near 0.52, roughly six times the registered 0.08 bar. The item-level calibration reported elsewhere in this document belongs to a later head and is a separate figure, from H20. That head was already well calibrated before any refit.
Repeating the same decision twenty times gave two different readings depending on how the test was run. In the first run a second job was sharing the endpoint. Jev changed its answer on 0.05 percent of repeats, while the open head changed on 4.95 to 5.85 percent across splits, so on that reading Jev looked like the steadier decider.
A follow-up test re-ran the open head with nothing else on the endpoint, replaying the same points once at one call in flight and once at four calls in flight. At one call in flight the open head changed its answer on 0.05 percent of repeats on both the onboarding and accounts-payable processes, which matches Jev's own figure exactly. At four calls in flight it rose to 3.4 percent on onboarding and 1.95 percent on accounts payable. Those disagreements clustered on a single worker thread, the signature of calls being batched together, which places the earlier gap in the serving stack. The shipped service now defaults to one call in flight, which costs throughput and leaves accuracy where it was. Where all twenty repeats agree, more than half of that agreement lands on the wrong answer, and that holds for both approaches.

A separate attempt tried to hedge the disagreement that remains by having the service abstain on the claims it is least sure of. The mechanically chosen band abstained on 1.8 percent of claims on one accounts-payable split and 0.7 percent on the other. At rates that low you can't separate the band's own effect from the ordinary noise of an independent rerun, so the test is reported as inconclusive. A cleaner version is already registered. It reads the band's effect from a single run's own trajectory instead of comparing two separate runs.
Another test was meant to rank Jev's chosen actions against the rule table's by replaying every alternative from a saved checkpoint. Because the rule table finishes every arm's branches, the test favors the rule table by construction, so it's being rebuilt and won't be reported as a ranking.
The evaluation set (eight hundred held-out cases, plus a two-hundred-case harder set with shifted formats and a policy change) was built and frozen before either run. An independent review checked the protocol three times before it was allowed to run, and once more before the retrained head's run. Afterward, a party who did not build the scoring code rescored both runs from the raw case files.
Carrying it to a second process
Everything above ran on one synthetic process, supplier onboarding. A follow-on test moved the same approaches, Jev included, onto a second process built the same way: accounts-payable exception routing.
On the new process's evaluation split the judgement head scored 0.650 and the rule table 0.648, level with each other, and the head's lead from the first process was gone. A learned head trained on the first process scored 0.522 there when pointed at the second one with no retraining. Trained instead on a hundred of the new process's own labelled cases, 271 decisions in all, it matched the judgement head. Budget for roughly a hundred labelled cases from your own process before you trust a learned head on it.
A follow-up check showed that the match holds without the learned head seeing the judgement head's own answers while training or while running, a question this brief had left open. A retrained version had those answers stripped out of its features and was served with no call to the judgement head at all, and it still matched. It scored 0.660 against the stacked version's 0.663 on one split, and 0.670 against 0.671 on the other, with both gaps well inside noise. On the larger split it also wrongly called a case done less often than every other approach.
Jev, tested on the same second process, came in at or below both the rule table and the judgement head. Its caution showed up as fewer wrong claims of a case being done and more unnecessary hand-offs than either, with accuracy at or below the other two.

This second process was built the same way as the first and shares its mix of case types. Whether any of this holds on a process built differently is still open, and that test comes next.
What it costs
The evaluation runs, on hardware the company already owns, cost nothing beyond electricity. A narrow, capped budget was set aside for two jobs that need a stronger outside model: consulting an expert model on the hardest cases, and a blind check of results by a model that never saw how any answer was produced. That phase hasn't run yet, and the sixty-dollar ceiling stays reserved until it does. Jev didn't need a reopened TypeSafe account. We reached it through OpenRouter at the $0.0000128 per question measured on 24 September, and that cost is already folded into the results reported above.
The limits of the claim
The claim here stops short of saying a company's own hardware beats a well-funded vendor's model at this task in general. It rests on a controlled comparison whose results point both ways. The first version of the open head didn't beat the rule table, and a retrained version did, by a wide margin, on two different test sets. On production readiness the claim is narrower than the accuracy numbers alone would suggest. The retrained head's wrong-claim rate first came in above the line this project set in advance. Three guards were tried for that failure mode, two of them model-based, and both of those failed their own test. A vote across several readings separated a wrong claim from a right one at an AUC of 0.56. A version that read the cited document itself did no better than a plain deterministic check, at an AUC of 0.52. None of the three is in use. The wrong-claim rate came down once we corrected the evidence pack, which the environment had been filling with supporting documents it never checked, and the corrected rate sits close to the rule table's own. We found that fix after the bar was first missed, by tracing why, and the results above say so. The claim does cover direction, now backed by a measured result as well as a prediction. A company can reproduce the pattern behind a typed-question decision service without a vendor and tune it on its own cases. It can measure the result in the open and run it on its own terms, provided the process around the model checks its own evidence too.
What a company will be able to do with the code
The service is a few hundred lines of Python with no third-party library required, and it runs in front of any model that exposes token-level probabilities, including one a company already has behind its own firewall. Once it is released, a team can point it at their own ollama or vLLM endpoint and write a handful of yes/no or multiple-choice questions about a case they already handle, and get calibrated answers, each with an audit trail. It is public on GitHub under Apache-2.0 (https://github.com/dtinholt/decision-head), after two independent security reviews, so a company's own team can read every line before trusting it with real cases.
The open-weights step: a first result
The first version of this step has now run. Alongside the retrained decision head above, a separately fine-tuned, smaller head was trained on the project's own labelled cases and registered as its own arm. It scored 0.701 on the fresh evaluation set, close to the retrained head's 0.728 without matching it. That gap doesn't clear this project's own bar for calling the two equivalent, so the fine-tuned version ships with an experimental label and stays out of the replacement role.
What the open alternatives showed
Two open encoders released by vendors, Laya and CLM-8B, appeared within days of this project's own head. Each answers a decision in one forward pass without reading the evidence behind a case. We ran both on this project's own benchmark instead of taking either vendor's numbers at face value.
Fine-tuned on 1,567 of our own labelled decisions, drawn from 363 onboarding cases, a small open encoder scored 0.815 on that process's harder set, above our judging head and well above the rule table's 0.565. The same encoder fell below the rule table it was supposed to beat when it had only 201 labelled decisions from 58 cases of a second process to learn from. Untuned, both encoders matched the rules and went no further.
For this encoder the size of its training set decided the result, since the smaller accounts-payable set left it with a habit that cost accuracy on that process.

We also checked calibration, since one vendor advertises it as a strength. The judging head's own probabilities were already well calibrated before this test ran. On the 400 shared cases the open encoder's calibration error was 0.0057 against the judging head's 0.0313, both small. The judging head reads the cited evidence and keeps the ledger, with an audit row behind each call, and an encoder does none of it.
The judging head stays the safe default. A learned mode gets promoted only after it passes a held-out check on your own data, whatever its vendor reports.