Building a Jev Lookalike in One Evening

I had been designing an experiment around a product called Jev, from a company called TypeSafe, and when I went to open an account, TypeSafe was not issuing new API keys at the time. I reached Jev through OpenRouter for the comparison instead.

Jev does something I think gets too little attention when people talk about AI in business. You give it the state of a case and a list of typed questions, yes-or-no or multiple-choice, and for each question it hands back a calibrated probability with no prose attached. Plenty of the decisions that sit inside a governed process suit that shape better than a chat response does. Does this case have the evidence it needs, does it need a document nobody has yet, should someone reason harder about what is already on file, is this case done? Each of those wants a probability and an audit trail, and a paragraph of prose gets in the way of both.

What building it myself buys, beyond spite

Being locked out pushed me toward a question worth asking whether or not TypeSafe ever reopens enrollment, which is what a company gets by running this in-house. I count nine things. Everything stays inside the network, so a pilot can start before a data-processing agreement is signed. The confidence numbers are calibrated on your own labelled cases, and you can hand the reliability table to your own auditor. Pointing the same service at a new process takes a question file and a calibration run, with no new contract to sign. Once you have banked enough of your own decisions, you can distil a smaller model trained on exactly those cases and measure what that buys you. Every answer carries an audit row holding the exact question, the raw probability, the model and the latency, so a decision from six months ago can be reconstructed in full. The same service runs unchanged on anything from a small local model up to a frontier model in your own cloud account. The whole premise is that each question gets answered in isolation, and the service tests that claim on your own data, so the evidence for it comes from your own cases. The running cost is whatever your own hardware already costs. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head, after two independent security reviews, so your security team gets to read every line before it touches a real case. A purpose-built model may still win at the job it was trained for, and running my own means I can measure the two side by side, which is what the rest of this post does.

I had already been reading two internal analyses of where Jev fits inside a governed workflow, along with a study my colleague Thordur Arnason had published on LinkedIn. He measured Jev against a frontier model at reviewing agent trajectories after the fact, on a benchmark called tau-squared-bench, and he reported what he found straight. Jev came close to Haiku 4.5 at flagging wrong database changes, four times faster, at roughly four percent of the cost, and open-weight models running on one of our own DGX Sparks scored no better than a coin flip at the same task. He was also open about a blind spot in his own study. An agent can fail by doing nothing wrong and still doing too little, and a review of recorded actions will never catch that, because there is no action to point at.

That gap, the question of what should happen next, was the one I wanted to test, and then the account page told me I couldn't.

The probe

Before writing the experiment off, I made one call to our own qwen3.8:27b model through Ollama's native chat route, with logprobs turned on, and asked it a single yes-or-no question. I wanted to see whether the model would hand back anything beyond a token. A probability came back in that one call. The model's belief about its answer was in the response all along, and most code that calls these models day to day ignores it.

A model already computes this number whenever it generates a token, so the experiment no longer hung on one vendor. Jev's trick is to ask one narrow question at a time and read the answer back as a probability, and anyone can reproduce that. TypeSafe trained its own model specifically for this. The pattern around the model needs a shared state, a single question per call, one token read as a distribution, and a calibration step fitted on your own cases, and any team can build it.

The service's own smoke test, run the same night, makes a better illustration than that single probe, because it asks all three question types about one realistic case at once. Treat it as an illustration only, since it was never meant as a benchmark. Below are the request and the response, with every number exactly as it came out of the run.

POST /v1/systemone
{
  "state": "Acme Logistics BV, a subsidiary of Acme Holding NV, submitted a
            certificate of insurance dated 2026-01-15 naming Acme Holding NV as
            the insured party rather than the subsidiary. No endorsement
            schedule or subsidiary rider was included.",
  "model": "jev-latest",
  "questions": {
    "insurance_satisfied": {
      "type": "noul",
      "instructions": "The evidence on file satisfies the insurance requirement for Acme Logistics BV."
    },
    "next_step": {
      "type": "choice",
      "instructions": "What should happen next?",
      "options": ["CONTINUE", "REQUEST", "HANDOFF"]
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgent is this case?",
      "levels": ["low", "medium", "high"]
    }
  }
}

200 OK
{
  "model": "qwen3.8:27b",
  "answers": {
    "insurance_satisfied": {"noul": 0.001},
    "next_step": {"choice": "REQUEST",
                  "probabilities": {"CONTINUE": 0.007, "REQUEST": 0.910, "HANDOFF": 0.083},
                  "confidence": 0.827},
    "urgency": {"score": "medium",
                "probabilities": {"low": 0.188, "medium": 0.412, "high": 0.400},
                "confidence": 0.012}
  }
}

Every number above comes from the nine calls that made up the smoke test. The service spotted the entity mismatch and said the certificate does not satisfy the requirement, with 99.9 percent probability. It judged the mismatch fixable and asked for the right document, keeping the case out of a person's queue, choosing REQUEST with probability 0.91 and a margin of 0.83 over the next option. On urgency it landed on medium by a hair, with low, medium and high almost evenly split and the model about as unsure as a model can be. An even split like that still carries information, because it says the model doesn't know, and calibration is there to keep that signal intact. The first call took 9.4 seconds, and a repeat with the same state already warm in the prefix took 2.7. To check the isolation claim directly, I asked every question a second time with the other two visible, and none of the three answers changed. The satisfied and request readings are logged in the fleet's governance log for 23 September 2026, wiki/governance/log-2026-09-23.md.

The model field in the request is an alias that lets a client keep its existing request exactly as it was, and the response always names the model that answered.

Three reviews, two refusals

Everything in our research process now goes through an adversarial review before it runs, and this design went through three rounds of it. The first round found nine problems. One was a scoring bug that quietly double-counted a budget. Another was a splits design in which two of the four outcome classes were nearly impossible to get wrong, which would have drained the test of any meaning without anyone noticing. The reviewer's verdict was that it must not run. We fixed each problem with its own test, rebuilt the evaluation set larger and harder, and sent the design back.

The second review caught something worse. The cheap mock worker we use for early testing had been marking a document as satisfying a requirement whenever it was present, without checking whether it named the right company or was still valid. With that shortcut in place, a wrong document on a real case would have gone straight through. We fixed the worker and spot-checked the real model against cases built specifically to expose that failure, and the model got every one of them right. The second reviewer still blocked the run.

The third review confirmed every fix independently and ran a statistical check against the evaluation data, which showed the design had enough statistical power to detect the effects it is looking for. It approved the run with conditions. We had to cap spending and add a shared rule module so two parts of the system can't quietly disagree about what "current" means. We also had to register the open decision head itself as a formal arm of the experiment, next to the rules-based approach, a free-choosing model, and the same local model asked Jev's exact questions.

A separate piece of work landed the same night. DeepSeek V4.1 Flash went into production on a shared endpoint that runs across both of our DGX Sparks as one model, serving 27 tokens a second on a single stream and 65 aggregate across six sessions, per the throughput ladder banked in experiments/dual-spark-tp2/results/verification-candidate2-2026-09-23.md. It became the assistants' own primary model, and it took over as the worker behind this experiment too. qwen3.8:27b, which answered the probe and the smoke test above, was the model I had on hand that night to prove the mechanism worked. The experiment itself runs on DeepSeek now, recorded as a fourth addendum to the registration.

The speed question

The open head takes about three seconds where Jev takes a third of one, and these are the measurements behind both figures.

Jev, through OpenRouter, answered a one-question probe in 0.35 seconds when I measured it on 24 September 2026. The call cost $0.0000128 for 304 input tokens.

On the same day, the open head's fast mode ran DeepSeek V4.1 Flash on our two DGX Sparks while four benchmark streams were already hammering the same hardware. PREREG-J6.md and results/design-review-J6-2026-09-24.md record these figures:

measure value
median 3.1 s
p90 4.8 s

No worst-case latency or request count was banked for this run.

Jev wins that comparison outright. Deliberate mode samples eight answers per call for the harder decisions, so it will cost roughly eight times as much per call, and I haven't measured its latency yet.

A gate like this fires once per checkpoint in a multi-step process. In the J5 eval2 run, 794 cases passed through 2,681 checkpoints, about 3.4 per case. The steps on either side of a checkpoint, somebody chasing a supplier for a document, say, or routing a case to an expert, take anywhere from minutes to days. Against a two-day wait for a document, three seconds won't register with anyone.

A gate's response time beside the steps around it
Seconds, logarithmic scale.

On 800 eval cases, the zero-shot open head scored 0.545 decision-class accuracy, against 0.575 for the rule table it's supposed to beat. The gap is -0.029, and the interval around it runs from -0.060 to +0.001, so I could not rule out that the rule table was better. Jev's own run landed the same way, at 0.540 on that eval set, 0.560 on dev and 0.536 on the 200-case set built with contradictory evidence, below the rule table on both, and this time with an interval that excludes zero. On the contradictory-evidence set, every arm, rules included, wrongly called a case complete 21 to 28 percent of the time. The test's own scoring caused most of those errors. The test sometimes plants a second, equally valid document and then marks a citation of it wrong. Once that was fixed, the rate dropped to between 1.5 and 9.5 percent depending on the arm.

Then I retrained the head and ran it again on 800 fresh cases. It scored 0.728 against 0.566 for the rule table, and 0.739 against 0.565 on the harder set, where it beat Jev's own 0.536 on the identical cases by 0.199, measured on the 199 cases both scored. Every one of those gaps excludes zero. It was the first result in the project where the open head beat both the rule table and Jev. On the same fresh set, though, the retrained head wrongly called a case complete at a rate of 0.047, which is above the line I set in advance as my trigger to redesign the safeguard meant to catch that mistake. I've built and tested three guards since, two of them model-based. The first looked perfect, catching 79 of 79, until a review found it had been scored against the same rule it was supposed to be checking, so it could never have missed. I threw that number out. A second, a vote across several readings of the case, separated wrong claims from right ones at an AUC of 0.56. A third read the cited document itself and did no better than a plain deterministic check. All three are retired, and the fix that worked went into the pack builder, which I cover under "Where this stands".

In return for the slower answer, the state of a live case stays inside our own network, and every answer above came with an audit row I can hand to someone six months from now. Retraining the head on a hundred of our own labelled cases doesn't need a GPU at all. Jev is faster and needs no infrastructure of its own to run, and it was tuned on more data than we'll log ourselves any time soon. The retrained head now beats the rule table on two test sets. Against Jev it won on the harder set, the only one where both ran in that round. Its wrong-claim rate was the problem still open after those runs, and the pack-builder fix later in this post is what dealt with it.

How consistent any of this is

I finally scored two questions I'd left blank since day one. One asks how often each approach disagrees with itself. The other asks whether its chosen move beats the alternatives once you replay the case forward. On the first pass Jev disagreed with itself on 0.05 percent of repeats and the head on 4.95 to 5.85 percent, which looked bad for the head. A second job of mine had been hitting the same endpoint during that pass, though, so I didn't trust the number. I re-ran the head alone on the same points with nothing else on the endpoint, first one call at a time and then four. With one call at a time, the head disagreed with itself on 0.05 percent of repeats, which matches Jev's own figure exactly. At four calls at a time, disagreement climbed to 3.4 percent, and the disagreements piled up on one worker thread, over and over. That clustering is the fingerprint of calls being batched together on the server. The inconsistency came from my serving setup, so the service now defaults to one call at a time, and the only price of that default is speed. The repeats also showed that where every repeat agrees, both approaches agree on the wrong answer more often than on the right one.

How often a repeated decision changes
Share of repeats that disagree with the majority answer, twenty repeats per decision point.

I also tried having the service abstain on the claims it is least sure of, in the hope that abstention alone would bring the wrong-claim rate down. It abstained so rarely that an ordinary rerun moved the numbers more than the abstaining did. I'm calling that test inconclusive. I have registered a cleaner version that reads the effect from a single run's own trajectory instead of comparing two separate runs against each other.

The replay question came back against Jev. The replay, though, finished every branch, the rule table's own included, by running the rule table forward, which stacks the deck before either side makes a move. I'm rebuilding that test and won't report it as a fair ranking.

Where this stands

The service is built and has now run twice, and both runs were rescored independently from the raw case files by someone who didn't build the scorer. It is a few hundred lines of Python that use nothing beyond the standard library. It runs against the DeepSeek endpoint on our own hardware, and the same code can point at a small model on a workstation or a frontier model in a cloud account. It calibrates on labelled cases you supply and tests its own claim that questions stay isolated from one another, and every answer gets an audit row.

The prediction I was least attached to came true. On the first evaluation set the open model asked Jev's exact questions scored 0.545 and the purpose-built Jev 0.540, and neither of them beat a plain rule table. Retraining the head on the project's own cases changed the answer. On a fresh eight hundred cases it beat the rule table by 0.161, and on the harder set it beat Jev by 0.199. I also ran a smaller, separately fine-tuned version of the head, the distillation step I said back then was still ahead of me. It scored 0.701 to the retrained head's 0.728, too far apart to call the two equivalent under the rule I set in advance, so it ships labelled experimental and the retrained head stays in place.

The retrained head's biggest number came with a cost. It wrongly called a case complete more often than the line I set for myself before I saw a single result. I built and tested three guards for that problem, two of them model-based, and none of them worked. I threw out the first because a review found it was scored against the same rule it was supposed to be checking. The second, a vote across several readings, reached an AUC of 0.56. The third opened the cited document and read it against the requirement, going further than any vote or date check, and it failed in the same way. Combined with a plain deterministic check, its numbers came out identical to that check running alone, and by itself it separated a wrong claim from a right one at an AUC of 0.52. No model-based guard added anything past the deterministic check. The pack builder now runs that same validity check, with a caveat. On my own test cases it catches wrong claims by the same rule my own generator used to label them wrong in the first place, so a real deployment would need a rule of its own.

Then I went looking for the reason the wrong-claim rate was high at all, so I could explain it as well as report it. Of the retrained head's 37 wrong claims on the fresh set, 31 were packs where the environment listed extra supporting documents it never checked, most of them decoys built to catch a careless citation and never meant to pass. Four were the worker citing the wrong document on a hard case, a mistake the plain rule table also makes on the identical cases, and the last two were further citation mismatches. An independent check traced none of the 37 to the head's judgement question or to its elimination walk. I fixed the pack builder to check every document the way the rest of the system already checks them, and re-ran the numbers:

Every case that changed moved from wrong to verified, and nothing else moved. Someone else checked all of it independently before I trusted it. The fixed reading meets the line I set for myself in advance and the original reading misses it, and I'm reporting both so the fixed one doesn't quietly replace the other. I also found the defect only after the line had first been missed.

Wrong-claim rate before and after the evidence-pack fix
Share of cases closed with a claim the registered scorer marks wrong.

Onto a second process

Every number above came from one synthetic process, supplier onboarding, so the obvious next step was to see whether any of it carries over to a different one. I ran the rule table, the judgement head and Jev on a second synthetic process built the same way, accounts-payable exception routing.

On the new process the rule table scored 0.648 on ap-eval and the judgement head 0.650, a gap well inside noise.

The head's gap to the rule table, run by run
Decision head minus rule table, decision-class accuracy. The first run used the zero-shot head.

Pointed at the new process with no retraining, the learned head dropped to 0.522 on ap-eval. Retrained on a hundred labelled cases from the new process itself, 271 decisions in all, it caught back up to the judgement head.

Jev scored below both the rules and the head on the new process. When I looked at the kind of mistakes it made, Jev claimed a case done incorrectly less often than either of the others and handed cases off unnecessarily more often, a cautious habit that cost it accuracy here.

Jev on the second process
Share of cases, Jev minus the judging head, with both sides scored on validated evidence packs. Left of zero means Jev does it less often.

A caveat in the first version of this section has since been settled. When I first wrote this up, the learned head's training features and its serving path both called the judgement head's own live output, so its parity with the judgement head could have been a classifier stacked on its own input. I registered a follow-up to check that directly. I stripped out every judgement-head feature, trained on the same hundred labelled cases, and served the result with no call to the judgement head at all, and it still matched. It scored 0.660 on ap-eval to the stacked version's 0.663, and 0.670 on ap-challenge to 0.671, both gaps well inside noise. On ap-eval it also claimed wrongly less often than every other arm. A hundred labelled cases and the structured features were enough by themselves, which retires the stacking caveat. The encoder I fine-tuned on the same split fell to 0.55, while the stacked head and this independent one both held at 0.66 to 0.67. The encoder's export keeps only the 201 decisions where something was still open, from 58 of those cases. The learned head saw those same 201 plus 70 closing checkpoints that label a case done. On those same open decisions the feature-based heads still beat the encoder, so in that comparison the feature design decided the outcome.

All of this rests on one new process built the same way as the first, and I still don't know whether it holds on a process built along different lines.

Two open models showed up, and I ran them anyway

Two free open encoders, Laya and CLM-8B, appeared within days of everything above. Both are both small, and neither reads evidence. I registered predictions before running either, as I do for everything else in this project, and I was wrong more often than I was right.

Fine-tuned on 1,567 labelled onboarding decisions, the open ones from the same training run my learned head used, Laya scored 0.815 on the harder set of the process it learned, against 0.739 for my judging head and 0.565 for the rule table. I had predicted near parity with the head and a loss to the rules, so both predictions missed in the same direction on the same split.

Accuracy against the rule table on the harder set of each process
Each dot is one approach's difference in decision-class accuracy from the rule table on the same cases. The right-hand column gives the accuracy itself. Fine-tuning and retraining used the 600-case onboarding training split and the 100-case accounts-payable development split.

Then I ran the same recipe on 201 labelled decisions from the accounts-payable process and expected the same story. Laya fell below the rule table it was supposed to beat. Across the encoder's own two training sets, with the model and the recipe held fixed, the result flipped, and in that comparison I put the flip down to the amount of labelled data. Two hundred decisions left it with a habit of promoting requests to consultations it didn't need to make, which it had not picked up from some 1,600.

Where the gains and losses sit, by decision class
Share of checkpoints of each true class that an approach classified correctly. The DONE class is left out because the harness decides it without a model call and every approach scores 1.000.

Untuned, both models sat level with the rules and moved nothing in either direction.

I also checked calibration, since Laya's maker leans hard on that number. The first version of the head, tested back in J1, was badly calibrated at the decision level, but at the item level my own head turned out to be well calibrated already, before I did anything about it. Both item-level errors were small, the open model's a little smaller than mine, and that edge was less than I'd expected going in.

The head keeps the job because it reads the document and writes a ledger row I can hand to someone later, which the open encoders can't do.