The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control

Dinand Tinholt Capgemini, Head of AI Center of Excellence, Americas 23 September 2026

This paper describes a registered experiment and the service it tests. Two rounds of it have run, and an independent party has rescored both from the raw case files. J1 was the original four-way comparison. J5, a second registration, retrained the head and tested it on a fresh split. Each number below comes from those two verified rounds, states a prediction and says so, or cites someone else's published work by name. Section 6 sets the measurements against the five pre-registered hypotheses and explains why the claim guard built for the retrained head's wrong-claim rate gave way to a fix in the evidence bookkeeping.

Abstract

Much enterprise AI work inside a governed process consists of decisions made at checkpoints. At each one the system has to pick among retrieving a document, asking a supplier for one, reasoning over what it already has, escalating to a stronger model, closing the case, or handing it to a person. TypeSafe's Jev is built for this kind of question. It receives a case and a state once, answers a set of typed questions about that state in parallel and in isolation, and returns a calibrated probability for each. Two outside studies have measured parts of this claim. One checked variance on five fixed responses; the other was a post-hoc audit that ranked Jev against a frontier model at flagging wrong actions already taken. Both scored answers to fixed or finished cases. This paper registers a test of the next-step choice made while a case is still open, with the model's stated probabilities scored against its own record. It also asks if the pattern needs a hosted vendor at all. The vendor closed enrollment on the day the test was designed, so the paper also builds and evaluates a self-hosted equivalent. This decision head reproduces Jev's request and response shape over any local model that exposes token-level probabilities. It makes one single-token call per question, calibrates on the operator's own cases and writes an audit row for every decision. We describe the mechanism, the experiment (six actions, four decision classes, five controllers, an obligation ledger that catches silent omission, and counterfactual replay from saved checkpoints), and the five pre-registered hypotheses with numeric predictions. On the first registered split the self-hosted head failed to beat a hand-written rule table, and so did the hosted Jev service measured afterward on the same split. A second registration retrained the head and tested it on a fresh split, where it beat the rule table by 0.161 to 0.171 decision-class accuracy and beat Jev by 0.199 on the 199 cases both scored, all three intervals excluding zero. Its wrong-claim rate on that split sat above the threshold this project set in advance to trigger a guard redesign. The redesign traced the defect to the workflow engine's own evidence bookkeeping. Once the pack builder validates every supporting document before a claim leaves, the retrained head claims wrongly at close to the rule table's own rate, and no model-based guard added anything beyond that check. We found the fix after the threshold was first missed and report it that way. A third registration tested transfer to a second synthetic process built with the same generator family. There the rule table matched the judging head and a zero-shot transfer of the learned head scored 0.522 against the judging head's 0.650; retraining on a hundred labelled cases from the new process recovered parity. Jev, the hosted reference, handed off more cases on the new process and claimed fewer wrongly, with accuracy at or below the rule table and the judging head. A party who did not build the scorer rescored all three results from the raw case files. A fourth registration ran two vendor-released open alternatives, Laya and CLM-8B, on the same benchmark. Fine-tuned on 1,567 labelled decisions from 363 onboarding cases, one encoder beat the retrained head by 0.07 to 0.08 and the rule table by 0.25 to 0.28. The same recipe fine-tuned on only 201 labelled decisions from 58 cases of the second process fell 0.08 to 0.11 below the rule table. Both fine-tunes claimed a case complete wrongly less often than the judging head, and one encoder's item-level calibration matched the head's own. Across the encoder's own two training sets the larger set won, and the learned mode should be checked on held-out data before it is promoted.

1. Introduction

A governed process depends on a small choice it makes over and over. Wherever the process could go one of several ways, something has to pick the way. In supplier onboarding or an accounts-payable exception queue, that choice seldom involves writing anything. The controller has to judge whether the evidence on file is sufficient, whether the case needs a document nobody has yet, whether interpreting what is on file needs a stronger model, or whether the case is done. Each wrong choice fails in its own way. Asking for more evidence when the answer was already on file wastes a cycle. Reasoning over a case that is missing a document produces a confident wrong answer. A compliance function worries most about a case declared complete while an obligation is still open, because nothing downstream catches it until an audit does.

TypeSafe built a product around a narrower, more testable form of this problem. Jev takes a state and a set of typed questions, each phrased as a yes/no ("Noul"), a choice among named options, or a score. For each question it returns a probability distribution, computed in parallel and in isolation, at $0.042 per million input tokens with output free (docs.typesafe.ai, quickstart and primitives pages; typesafe.ai/blog/introducing-system-one-models-and-jev). The vendor's own evaluations page reports 67.8 percent average accuracy at $0.0004 per case and 0.4 seconds, self-reported, with a reference-model bias the vendor itself acknowledges (evals.typesafe.ai). Two outside parties have since tested parts of the claim independently. A LangChain post dated 20 September 2026 ran five fixed weather-agent responses through Jev one hundred times each and found the oracle answer reproduced on all five hundred calls, with per-case variance of 1.5e-5, at $0.00035 per call and 0.44 seconds. That run measured variance on five items whose answers were known in advance. Thordur Arnason's LinkedIn study, dated 23 September 2026, took completed agent trajectories on τ²-bench (320 conversations, 470 recorded database changes, 22 of them wrong) and asked whether Jev could rank the wrong changes for a human reviewer better than a frontier model. Jev scored 0.74 against Haiku 4.5's 0.77 (0.5 is chance) while running about four times faster at about four percent of the cost, and code-plus-narrow-questions doubled the catch rate inside the top ten percent of a review budget. He also ran the open-weight models Laya and Nemotron 3.5 Lightning on a single DGX Spark, and both ranked the same residual errors at about chance.

Neither study tested the question a governed process needs answered most. Both reviewed actions after the fact. A controller has to decide what should happen next before any action is taken, while the case is still open and a wrong choice has time to compound. Thordur named the gap in his own write-up. In 31 of his 53 failed conversations the agent made no wrong database change at all; it failed by doing too little. A reviewer that only ranks recorded actions has nothing to rank in such a case. He also left unmeasured whether Jev's stated confidence tracks its own correctness. His stated next step is local fine-tuning, which raises a separate question this paper answers first, namely whether the value sits in the hosted model or in the shape of asking. If a local model given identical typed questions in an identical format performs close to Jev, the pattern carries the value and can run anywhere a model with token probabilities can be reached. If the local model falls well short, the hosted model is doing something local models cannot yet do.

We register a prospective, next-step decision test with four things neither prior study measured: whether a controller can tell "the case needs a document that is not on file" from "the case needs interpretation of a document already on file," whether acting on that distinction completes more cases correctly at lower cost, whether the controller's stated probability of being right tracks being right, and whether an open, self-hosted model asked the same typed questions gets most of the way to a hosted one. TypeSafe closed enrollment to new API keys on the day this protocol was written, so to answer the fourth question we built a small service, described in Section 3, that answers Jev-shaped questions over a model we run ourselves.

Jev and the System One family. TypeSafe describes Jev as a "System One" model, a fast, narrow model that answers typed questions, set against slower models built for open reasoning (typesafe.ai/blog/introducing-system-one-models-and-jev). The vendor documents nine known failure modes for the current version, jev-1.13.0, including literal reading, arithmetic, dates, indirection, large irrelevant state, adversarial content, contradictory criteria, structural invariants, and generation (docs.typesafe.ai, model-jaggedness/jev-1.13). Few vendors publish where their own product breaks. This documentation shaped several design choices in Section 5, and it is why code handles arithmetic and date logic everywhere in this experiment.

Routing. A separate line of work treats model choice itself as a decision problem, in which a query goes to a cheap model or an expensive one depending on predicted difficulty. RouteLLM and similar frameworks make this choice with a trained router placed in front of the models it routes between. The decision head in this paper is related and narrower. It answers a small set of fixed, typed questions about a state in order to choose the next action in a workflow, and several of those actions, such as retrieving a document or closing a case, involve no model call.

Value of information and selective prediction. The INFO-versus-REASON distinction in this paper comes from decision theory. Before spending reasoning effort, the controller asks whether the missing piece is a fact that could be fetched or an interpretation that has to be made. Selective prediction lets a system abstain when its own confidence is low. The calibration literature measures whether a stated probability tracks empirical correctness. Both bear on whether a typed-question controller's probabilities are usable for anything beyond picking a top choice. Expected calibration error, the metric H3 is scored on, comes from the calibration literature.

τ²-bench and post-hoc review. Thordur Arnason's study, summarized above, is the closest existing measurement of Jev's practical value and the direct predecessor to this registration. He ranked a finished trajectory after the fact, and this paper's design is built around a prospective next-step decision on an open case.

LLM-as-judge consistency. Arm D in this design uses a frontier model as an expert and, in phase J3, as a blind adjudicator over a holdout sample. LLM judges are known to give inconsistent answers across repeated calls, orderings and phrasing, so the adjudication sample is drawn once from a fixed seed and scored blind, with no reruns in search of a preferred answer.

Open single-pass alternatives. Two open decision models appeared within days of each other in late September 2026, described here as their vendors describe them in syn-open-decision-models-2026-09-28, dated 28 September 2026. Laya, from Convai Innovations, released 18 to 21 September 2026 under Apache 2.0, is a 421-million-parameter ModernBERT-large encoder that answers typed choice, score, and yes/no primitives in a single forward pass, with a per-question-type temperature refit the vendor reports moving expected calibration error from 0.21 to 0.47 raw down to 0.08 to 0.11. CLM-8B, from Stanford and NVIDIA, released 23 September 2026 under Apache 2.0, is a frozen Qwen3-8B encoder behind two small contrastive heads that rank a given set of candidate actions by cosine similarity, and by the vendor's own account it has no calibration step. Both are single-pass encoders over typed primitives, and neither reads the evidence documents behind a case. Until now every published comparison between either model and Jev has come from a vendor or the press. Section 6.7 registers and reports the first run of both models on this project's own pre-registered benchmark.

3. The decision head

Mechanism. A decision head takes one state and a set of typed questions about that state and returns one probability distribution per question. Three question types are supported: a Noul, a yes/no probability; a Choice, a probability over a small named set of options; and a Score, a probability over an ordered set of levels. The state is rendered once into a shared prefix (a system turn plus a user turn holding the state text), and every question is then asked as its own separate final turn appended to that same prefix, in its own model call. No question's prompt contains another question's text. That arrangement is what this implementation means by "evaluated in parallel and in isolation." The same arrangement lets a serving engine's prefix cache compute the shared prefix once, so the second and later questions in a batch cost less than the first.

From state to audited answer
Figure 1. From state to audited answer.

Each call asks the model for exactly one token, at temperature zero with logprobs turned on. The service reads the response's top-token log-probabilities and normalizes them over the tokens the question accepts ("Yes" and "No" for a Noul; a letter per option for a Choice or Score). Probability mass that fell on any other token is reported as unparsed mass. If none of the accepted tokens appear at all, the answer degrades to a flagged uniform distribution, so the service never fabricates a number. The read is mechanical. It takes a distribution the model already computes at that position, with no second model call and no sampling.

The service tests isolation by answering the same question set twice. The first pass uses the normal isolated form. The second adds every other question's instructions to the shared prefix, so the model sees the full list while still being asked for one token at a time. The two runs are compared per question, reporting the raw probability drift and whether the winning answer changed. A client can therefore check a vendor's isolation claim as a number on their own state.

Calibration applies temperature scaling to the logit of a probability: p' = sigmoid(logit(p) / T), with the temperature fitted per question type on a labelled development set and applied at answer time. For a Noul question this is a direct transform of the single yes-probability. For a Choice or Score question, where more than one probability exists per question, the same transform is applied to every option and the result is renormalized to sum to one; this degenerates correctly to the Noul case when there happen to be exactly two options. Confidence for a Choice or Score answer is reported as the top probability minus the second, computed after calibration so the reported confidence always matches the reported probabilities.

Every decision produces an audit row holding the question asked, the raw and calibrated probabilities, the winning answer, latency, and a hash of the exact request. The log holds only what was sent to the model, and secrets stay out of every log line. Where an API key is needed, the service reads it from a named environment variable at call time, and the key appears nowhere in code, configuration or output.

A worked example, for illustration only. The service's own smoke test (tools/decision-head, nine calls against qwen3.8:27b, the model used for this validation run) puts a real state through all three question types at once. State: "Acme Logistics BV, a subsidiary of Acme Holding NV, submitted a certificate of insurance dated 2026-01-15 naming Acme Holding NV as the insured party rather than the subsidiary. No endorsement schedule or subsidiary rider was included." A Noul asking whether the evidence on file satisfies the insurance requirement for Acme Logistics BV answered 0.001, so the model read the entity mismatch and said no with near certainty. A Choice over the next step (CONTINUE, REQUEST, HANDOFF) answered REQUEST at 0.910 against 0.007 for CONTINUE and 0.083 for HANDOFF, a margin of 0.827 over the next option. The certificate names the wrong entity, and the parent-subsidiary relationship makes this a gap the supplier can likely close. A three-level urgency Score answered medium by a small margin, at 0.188 low, 0.412 medium, 0.400 high, a margin of 0.012. The near tie shows the model uncertain between two adjacent levels, which is what a calibration step should surface. The three questions took 9.4 seconds on the first call and 2.7 seconds once the shared prefix was warm, and the isolation check (the alone-versus-together comparison above) changed none of the three decisions. qwen3.8:27b was the model available for this validation run and for the earlier single-token probe that first showed a logprobs read was possible; Section 5 describes the different, larger model the registered experiment runs its worker and arm C′ against. The satisfied and REQUEST readings above are recorded in the fleet's governance log for 23 September 2026, wiki/governance/log-2026-09-23.md.

The API shape. The service exposes POST /v1/systemone, taking a state (string or object), a model field (read but not routed on: a client written against the hosted Jev API can point at this server by changing only the base URL), and a map of typed questions, and returning a matching map of answers, a usage block, and the model id the server was started with. A GET /health endpoint reports backend and model; GET /v1/models reports the real model id and any aliases a caller might send. The request's model field is advisory so that a hosted call can be replaced by changing one URL, with no code rewritten.

Backends. Two real backends are implemented, sharing one interface (token_distribution(prefix, question, allowed_tokens) -> distribution). Ollama is called on its native /api/chat route only, with think:false. On the same server the OpenAI-compatible /v1 route passes through a thinking wrapper, and on the fleet's own host it returned the wrapper's filler token in place of an answer, so the service calls only the native route. An OpenAI-compatible backend covers vLLM-style servers and turns off template-level thinking through chat_template_kwargs.enable_thinking:false. Arm C′ calls this backend in the registered experiment, against the fleet's own dual-Spark endpoint (Section 5). A third backend generates deterministic fake distributions for tests and needs no model at all. The whole service uses the Python standard library with no third-party dependency, so it runs unmodified on a DGX Spark, on a workstation GPU, or on a small server in front of a hosted gateway.

4. What a self-hosted decision head adds and what it costs

The case for a head built on the operator's own model rests on nine properties, each traceable to a design choice in Section 3, and it makes no performance claim against a hosted model.

Data stays on the network. The state, the documents behind it, and every answer stay on hardware the operator controls, because the service and the model it calls both run there. A pilot can start before the data-processing agreement and the vendor security questionnaire are finished, since no vendor boundary exists for that paperwork to describe.

Calibration the operator owns. The temperature-scaling fit in Section 3 runs on the operator's own labelled cases and produces a reliability table the operator can hand to their own auditor. A hosted service sets its stated confidence through its own process on its own cases, and the operator cannot refit it when their documents change.

Any question set, any process. A new use case needs a new question file and a calibration run against a labelled sample. The isolation and audit mechanisms in Section 3 apply to any typed question, so one head can be pointed at a second process in about a day.

Fine-tuning, with a measured transfer cost. Once enough operator-specific decisions have accumulated, a smaller model can be distilled on exactly those cases. That step is registered separately as phase J4, with its own pre-registration and adversarial review; a fine-tuned learned head, arm L, first ran in J5 (Section 6.3). Section 6.6 measures what a change of process costs, where a hundred labelled cases from the new process recovered the parity zero-shot transfer could not reach.

Audit per decision. Every answer carries a row with the raw distribution, the exact model id, the question-set version, and the latency (Section 3). Someone reviewing a decision six months later can rebuild from that row alone what the model saw and returned, along with what a downstream system did with it.

Choice of base model. The same service runs unmodified over the model this experiment now uses (DeepSeek V4.1 Flash, served as one model across two GPUs, Section 5), over a small model on one workstation, or over a frontier model behind the operator's own gateway. Only the backend configuration in Section 3 changes.

Measured isolation. Section 3's isolation check answers a question set twice, once in isolation and once with every other question visible, and reports where the two disagree. The operator can produce that number on their own cases.

Cost from the power bill. The service has no per-token charge and no signup queue, and it costs whatever the hardware already costs to run. The model file is pinned by hash, so no vendor can repoint an alias without notice.

Open source. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head, release v0.1.0, after two independent security reviews (23 September and 6 October 2026), and an operator's own reviewers can read every line that touches their data before a single case is sent through it.

All of this has a price, and a hosted, purpose-built model keeps advantages a self-hosted head does not inherit. jev-1.13.0 was trained for typed decisions. The decision head reads one token at a time from a general model never trained for that framing, and Section 5's H1 measures how much that gap costs. TypeSafe publishes a self-reported accuracy figure across four workflows on its own evaluation set. The decision head in this paper had no published track record beyond the pre-registered numbers its results section fills in. A hosted service also comes with a vendor SLA and a support line, while a self-hosted one gets whatever support the operator's own team is prepared to run. This paper takes no side on these differences. Section 5's experiment measures those that synthetic cases can measure today.

What an enterprise gate should optimise for

For most enterprise workflow gates the question that matters is whether the gate chose right, and the measurements below put speed and correctness side by side.

Jev, reached through OpenRouter, answered a one-question probe in 0.35 seconds. The call cost $0.0000128 for 304 input tokens. Both measured 24 September 2026.

The open head's fast mode, DeepSeek V4.1 Flash served across two DGX Sparks with four benchmark streams already contending for the same hardware, answered a yes/no question at the following latency. These figures come from PREREG-J6.md and results/design-review-J6-2026-09-24.md. No maximum latency or request count was banked for this run.

measure value
median latency 3.1 s
p90 latency 4.8 s

Jev answered faster on both of those numbers. Deliberate mode, which samples eight answers per decision for harder cases, is expected to cost roughly eight times a single fast call; its own latency and accuracy have not been measured yet.

The latency numbers leave out the job a workflow gate does. The J5 eval2 run logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case before a case closes (results/verification-J5-2026-09-26.md). The steps a checkpoint gates, such as retrieving a document or routing a case to an expert, run minutes to days. A person waiting on a two-day retrieval step loses almost nothing to a three-second decision inside it.

Correctness is where the arms differ. On the first registered split, neither the zero-shot open head nor the hosted Jev service beat the hand-written rule table; Section 6.1 reports the full comparison. On that same split's hardest cases, every arm appeared to claim a case complete wrongly in 21 to 28 percent of cases. Most of that traced to one scorer defect, and Section 6.4 reports the flawed reading beside the corrected one.

A second registration retrained the open head and tested it on a fresh split, where it beat the rule table and Jev with every interval excluding zero (Section 6.3). Its wrong-claim rate sat above this project's own pre-registered trigger for a guard redesign, and the cause turned out to lie in the workflow engine's own evidence bookkeeping (Section 6.4). This paper reports the accuracy gain and the wrong-claim problem together, since either one alone would misstate what the project has shown.

The head was designed with latency as a secondary concern. A self-hosted gate offers accuracy measured on the company's own process, and a live workflow's state stays inside the company's network. The model can be retrained on a hundred of the company's own labelled cases without a GPU. Every decision has an audit row behind it, and changing behaviour means relabelling data where a hosted service would need a vendor ticket. A hosted service answers in 0.35 seconds with no infrastructure to run, and it is tuned on far more data than any one company will log on its own. Those advantages are real, and they make Jev a serious reference for this project to measure against. For the gate this paper describes, adoption should turn first on accuracy and then on who owns the data and what it costs to change behaviour.

5. The workflow-completion experiment

Task and environment. The controller under test sits inside a synthetic supplier-onboarding workflow. A local AI worker has already extracted findings from a case; the controller's job is to choose the next of six approved actions: CONTINUE, RETRIEVE(category), REQUEST(item_id), CONSULT_EXPERT, HANDOFF(item_id, issue), or CLAIM_COMPLETE(pack). Each case carries a policy version, an application, a set of requirements, documents already on file, a hidden internal store retrievable by category, what a supplier could additionally supply if asked, and a hidden key describing, for each open requirement, which of four resolution classes it belongs to. Every controller gets the same fixed budget per case of twelve steps, two expert calls and three requests.

The four-way class. Every open item in a case needs exactly one of four things. Some items lack evidence; RETRIEVE resolves them if the evidence exists internally, and REQUEST if only the supplier can provide it. Other items need interpretation of evidence already on file, handled by CONTINUE when the local worker can make that judgment and by CONSULT_EXPERT when the case needs a stronger model. When every obligation is closed, the case needs CLAIM_COMPLETE. An item that no one can supply goes by HANDOFF to a person, with the issue named. The design names the two errors it is built to catch. info_as_reason spends reasoning effort on a case that was missing a document. reason_as_info asks for a document when interpreting the existing evidence would have covered the requirement.

What a gate decides at each checkpoint
Figure 2. What a gate decides at each checkpoint.

Arms. Five controllers are compared. Arm A is a rule table over the worker's findings and the obligation ledger, calling no model. Arm B is the worker's own model choosing freely among the six actions from a single JSON reply. Arm B′ asks that same model the identical typed questions the decision head asks, with probabilities estimated from eight samples at temperature one where the endpoint does not expose logprobs directly, mapped to an action through the same fixed decision table used downstream. Arm C′, the decision head described in Section 3, asks the same questions over the fleet's standing endpoint, one single-token call per question. The worker behind arms A, B, B′, and C′ is DeepSeek V4.1 Flash, a three-bit-per-weight EXL3 quantization served as one model across both of the fleet's DGX Sparks, reached over the OpenAI-compatible route described in Section 3. The original registration planned qwen3.8:27b as the worker, and PREREG-J1 addendum 4 records the change. The 27-billion-parameter model's role in this project ends at the single-token probe that first showed a logprobs read was possible and at the service's own validation smoke (Section 3); J1 measures a different model. Arm C, the hosted Jev service, ran through OpenRouter from J2 onward. A fifth arm, D, hands the entire workflow to a Bedrock model with no typed-question structure, as a cloud-first reference point with no causal role, and runs only in the credits-funded phase described below. A selftest asserts that B′ and C′ share byte-identical state compaction and question text, so any gap between them comes from the model.

The obligation ledger and counterfactual replay. At the start of each case, one obligation is opened per requirement, closed only by a finding of met. A CLAIM_COMPLETE issued while any obligation is still open is recorded as a premature claim the instant it happens, whether or not a downstream guard later catches it. The ledger exists for the failure mode Thordur Arnason's study could not see, an agent that makes no wrong move and still does too little. Here that agent leaves an open obligation with no action taken against it, and the ledger records the gap even if no wrong action ever occurred. Every checkpoint can also be replayed. The environment snapshots its state before a decision and can restore it, so each of the six legal actions can be tried from the same starting point and scored against the same key. The resulting regret figure (best available outcome minus the outcome of the chosen action) separates the cost of picking wrong from noise in how the episode happened to unfold afterward.

Metrics. The primary metric is four-way decision-class accuracy; exact-action accuracy is reported alongside as secondary. info_as_reason and reason_as_info are reported as rates on the checkpoints where each applies. Per case: verified completion (claimed, pack matches the key, every obligation closed), premature-claim rate, unnecessary handoffs on cases that were resolvable, correct blocks, expert calls, requests, steps, cost, and latency. For controllers that report a probability, calibration is scored with Brier score, expected calibration error over ten bins, a reliability table, and a selective-prediction curve showing decision error against coverage as the controller abstains below rising confidence thresholds. Consistency is measured on one hundred fixed checkpoints replayed twenty times each, as per-decision agreement and, separately, as the share of decisions that are consistently wrong.

Hypotheses and predictions. Five hypotheses are pre-registered, each with a numeric prediction stated before any eval-split data exists.

H1, discrimination. Arm C′ classifies the next-step decision more accurately than arm A and than arm B′, tested on the four-way class with a paired, case-clustered bootstrap confidence interval on the 800-case eval split. Predicted eval-split class accuracy: A, 0.62 to 0.74, point estimate 0.68; B, 0.50 to 0.66, point estimate 0.58; B′, 0.60 to 0.74, point estimate 0.67; C′, predicted between B′ and the hosted arm on every row, class accuracy 0.64 to 0.76, point estimate 0.70 (the hosted arm C, where it runs, is predicted at 0.68 to 0.82, point estimate 0.75).

H2, completion at cost. Acting on C′'s choice yields more verified completions and fewer premature claims per dollar than arm A, at equal or fewer expert calls. Predicted: premature-claim rate 0.06 for A and roughly the same or better for C′; verified completion 0.55 for A against a predicted improvement for C′; expert calls per case roughly 0.9 for A against a predicted reduction for C′.

H3, calibration. C′'s probabilities predict its own correctness (expected calibration error at or below 0.08 after temperature scaling fitted on the development split and reported on eval) in a way the local model's own sampled probabilities in arm B′ do not (expected calibration error at or above 0.15). Before temperature scaling, C′'s expected calibration error is predicted between 0.08 and 0.15. One structural limit is registered in advance. A correctly timed CLAIM_COMPLETE closes every obligation and therefore has no open item and no corresponding question, so calibration is scored on INFO, REASON, and BLOCKED decisions and excludes true-positive DONE decisions.

H4, consistency at scale. On the fixed hundred-checkpoint, twenty-repeat set, Jev disagrees with itself on fewer than two percent of decisions, reported alongside the share of decisions that are consistently wrong. C′ is registered to sit between the typed-sample arm and Jev on this measure; the two-percent line applies to Jev alone.

H5, counterfactual value. Replaying every legal action from saved checkpoints, C′'s chosen action carries lower mean regret than arm A's, with the gap concentrated in the INFO/REASON confusion and away from DONE/BLOCKED decisions, tested with the same case-clustered bootstrap as H1.

Decision rules. H1 is supported if C′'s class accuracy exceeds both A's and B′'s with a confidence interval excluding zero. H2 is supported if C′'s verified-completion rate is at least A's and its premature-claim rate is lower, both with a confidence interval excluding zero, at expert calls no higher than A's. H3 is judged against the expected-calibration-error thresholds above; H5 is supported if C′'s mean regret is lower than A's with a confidence interval excluding zero. Two negative results are registered in advance as informative. A B′ that matches C′ would place the value in the typed-question pattern itself, which can run on commodity hardware. An A that matches C′ would mean typed questions add nothing a rule table lacked, and the vendor's own design documentation says such a result should be taken seriously.

Phases and cost. J0, harness and generator selftests plus a smoke run with a deterministic mock worker, ran on a laptop at zero cost. J1, arms A, B, B′, and C′ with the real worker (the dual-Spark DeepSeek V4.1 Flash endpoint, PREREG-J1 addendum 4) across the dev, eval, and challenge splits, plus the consistency and replay runs, ran on the fleet's own Sparks at zero marginal cost. J2, the hosted Jev arm, ran through OpenRouter, with a probe call confirming the request shape before the full run, at a projected cost near one dollar. J3 covers expert consultation through a Bedrock model across every arm, reference arm D, and a sixty-case blind semantic adjudication sample. It runs under a fixed credits ceiling of sixty dollars, and an explicit rule limits the credits to judging and evaluation. They never pay for generating training content. The same constraint governed how this project's synthetic data was built, since case generation is template-based and deterministic from a seed, with no model involved.

Adversarial review. The registration cleared only after three independent adversarial reviews, and the first two returned must not run. Along the way the reviews found and fixed a budget-accounting bug and a decision path that let CONSULT_EXPERT collapse onto the same underlying probability as CONTINUE. They also found splits where two of the four resolution classes were nearly impossible to get wrong. A mock worker had made two harder resolution classes unloseable by marking a document as satisfying a requirement whenever it was present at all, without checking entity or currency. Each was corrected and verified, with the splits rebuilt to eight hundred eval cases and two hundred challenge cases and every blocker tagged with the failure mode it exercises. The third review reproduced every prior fix independently, ran a Monte Carlo simulation against the actual eval-case distribution reaching roughly ninety-nine point eight percent power at the registered effect sizes, as PREREG-J1 addendum 3 records, and returned may run, with conditions: a cap of six thousand calls and five dollars for the decision head's own arm, a registered two-hundred-case Bedrock subsample inside the J3 ceiling, and one shared rule module, used by both the case generator and the mock worker, so that document currency and entity matching are judged identically everywhere they are asked. Arm C′, the self-hosted decision head, was registered in this same review, on the day TypeSafe stopped issuing new keys.

What ran after J1. J1 answered H1 through H5 on the arms described above. Two follow-on registrations are reported alongside it in Section 6. J5 retrained the decision head on the same task, calling the result arm C″, added a separately fine-tuned "learned" head, arm L, and tested both on a fresh 800-case eval split (eval2) and the existing 200-case challenge split, under the same adversarial-review discipline as J1: a design review before any run, and an independent rescoring from the raw case files afterward. J6 and J7 registered a claim guard meant to catch the wrong-claim problem the challenge split exposed. J6's design failed its own pre-registered gate, and J7's first result was withdrawn after review found it circular. Section 6.5 reports why the guard was retired in favour of a fix to the pack builder.

6. Results

J1 and J5 have run and were independently rescored from the raw case files by a party who did not build the scorer; every cell checked in that rescoring reproduced the launcher's own numbers within tolerance. The tables below report what was measured. J3, the expert-consultation and blind-adjudication phase, has not run and is reported below as an open item. The claim guard registered in J6 and J7 was retired (Section 6.5).

6.1 J1: the first registered comparison

Table 1. Decision-class accuracy, eval split (n = 800 cases).

measure A (rules) B (free LLM) B′ (typed, 8-sample, n = 400) C′ (zero-shot decision head) C (Jev, hosted)
decision-class accuracy 0.575 0.527 0.583 0.545 0.540

C′ − A: −0.029, 95 percent CI [−0.060, +0.001], interval includes zero. C − A: −0.034, 95 percent CI excludes zero. Jev also ran on dev (0.560) and on the 200-case challenge split (0.536); C − A on challenge is −0.042, 95 percent CI excludes zero. H1 is rejected. Neither the zero-shot decision head nor the hosted service it was built to match beat the rule table on this split, and in the two comparisons where the interval excludes zero (C versus A, both splits), it lies on the side opposite the hypothesis.

H2 asked whether acting on C′'s choice yields more verified completions and fewer premature claims per dollar than the rule table, at equal or fewer expert calls. H2 is rejected. C′'s verified completion falls below A's on eval by −0.126, with a confidence interval excluding zero, and on challenge by −0.060, again with an interval excluding zero. C′'s premature-claim rate on eval is higher than A's by +0.004, with an interval of 0 to +0.009. C′ uses roughly a hundred times A's expert calls on eval, 0.376 against 0.004 per case, and ninety-three times on challenge, against a registered prediction of equal or fewer. results/verification-J1-2026-09-24.md, lines 181 to 188, verifies these figures.

H3 asked whether C′'s probabilities predict its own correctness, with an expected calibration error at or below 0.08. H3 is rejected for C′. Its expected calibration error is 0.5247 on eval and 0.5312 on challenge, more than six times the registered ceiling. B′'s sampled probabilities score an expected calibration error of 0.6286 on eval and 0.6265 on challenge, which confirms the half of H3 that predicted the local model would be uncalibrated. The Jev-specific reading of H3, on the hosted arm C, was never run. These figures are verified in results/verification-J1-2026-09-24.md, lines 190 to 197.

6.2 Consistency, replay regret, and the abstention band: H4, H5, H19, and H21

H4 asked whether Jev disagrees with itself on fewer than two percent of repeated decisions, replayed twenty times each on the same hundred fixed checkpoints per split. The two-percent line was written for Jev, the hosted reference, and Jev meets it. The open head was registered only to sit between the typed-sample arm and Jev, which it does. Jev disagrees with itself on 0.05 percent of repeats on both the eval and the challenge split. The open head disagrees on 4.95 percent of repeats on eval and 5.85 percent on challenge, inside the wider band an earlier addendum predicted for it and above the two-percent line. These J1 repeats ran at parallel two on an endpoint that a second campaign's accounts-payable run was sharing at the same time; the re-measurement below, under H19, shows what that sharing cost. Counted per decision point, Jev's twenty repeats are unanimous on 99 percent of points on both splits. The open head's twenty repeats are unanimous on 82 percent of eval points and 71 percent of challenge points, so on this reading Jev decides more consistently.

The companion statistic registered alongside H4 measures whether those consistent answers are right. Of the points where all twenty of Jev's repeats agree, 65 percent on eval and 60 percent on challenge agree on the wrong action. For the open head the figures are 54 percent on eval and 51 percent on challenge. Where either controller answers consistently, the shared answer is wrong on more than half of those points.

H5 asked whether Jev's chosen action carries lower replay regret than the rule table's, with a confidence interval excluding zero, where regret is the value of the best available action at a checkpoint minus the value of the action taken. Replaying every legal action from every saved checkpoint and scoring a case-clustered paired bootstrap shows Jev with the higher regret.

Table 2. Mean replay regret by arm, case-clustered paired bootstrap, 10,000 resamples, seed 0.

split A (rules) C′ (open head) C (Jev) C − A C′ − A C′ − C
eval 0.013 0.115 0.069 +0.054 [+0.042, +0.066] +0.082 [+0.069, +0.096] +0.028 [+0.015, +0.042]
challenge 0.111 0.196 0.177 +0.047 [+0.004, +0.093] +0.052 [+0.022, +0.082] +0.004 [−0.042, +0.049]

Both intervals for C against A exclude zero, and both sit on the side opposite the registered prediction, with Jev's chosen action carrying significantly higher regret than the rule table's. H5 as registered is refuted.

The instrument behind this comparison has a structural bias toward the rule table, separate from any scoring error. The deterministic rules controller completes every branch replayed from a checkpoint, including the branch the controller under test chose. For arm A, its own chosen continuation and the yardstick it is measured against are therefore the same policy from the second action onward. The share of checkpoints where the chosen action already equals the best-value action shows the effect, at 98.7 percent on eval and 88.9 percent on challenge for the rule table, against 93.1 and 82.5 percent for Jev, and 79.6 and 68.7 percent for the open head. Part of the rule table's low regret is therefore built into the yardstick. The refutation stands as arithmetic under the registered instrument, and J7 registers a replay in which the arm under test supplies the continuation policy in every branch, with no prediction carried over from H5.

J7's H19 re-measured the open head's consistency under conditions J1 did not have: the same hundred checkpoints per split, replayed twenty times each, on an endpoint with no other campaign attached. Each split was run twice, once at one call in flight and once at four calls in flight (PREREG-J7 addenda 4, 5, and 7). At one call in flight the open head disagrees with itself on 0.05 percent of repeats on both the onboarding eval split and the new ap-eval split. That figure equals Jev's own J1 figure exactly. At four calls in flight the open head disagrees on 3.40 percent of repeats on onboarding and 1.95 percent on ap-eval. Table 3 reports all four cells, independently recomputed from the raw repeat files against the launcher's own numbers (PREREG-J7 addendum 8; results/verification-J7-H19-H21- 2026-09-29.md).

Table 3. J7 serving-consistency re-measurement, arm Cprime, 100 points times 20 repeats per cell, the same 100 points in both conditions within a split.

split calls in flight mean agreement per-repeat disagreement
onboarding (eval) 1 0.9995 0.05%
onboarding (eval) 4 0.966 3.40%
ap-eval 1 0.9995 0.05%
ap-eval 4 0.981 1.95%
How often a repeated decision changes
Figure 3. How often a repeated decision changes. Share of repeats that disagree with the majority answer, twenty repeats per decision point.

At four calls in flight the disagreeing repeats cluster on a single worker-thread residue class among the twenty draws. A thread pool batching calls together leaves that pattern, whereas a model answering differently each time would spread the disagreements evenly. Both concurrency-four figures sit below J1's per-repeat 4.95 percent on eval and 5.85 percent on challenge, measured under the shared endpoint. The inconsistency measured in J1 came from the serving stack. Ambient load from a second campaign sharing the endpoint caused it first, and batching added to it once concurrency rose. The shipped service now defaults to one call in flight. That default costs throughput, since the service serialises calls that could otherwise run in parallel, and leaves accuracy unchanged.

H19's ap-eval condition required a fresh source run on arm Cprime that does not otherwise exist in this project. Scored on its own, that run reaches 0.533 decision-class accuracy on ap-eval against 0.648 for the rule table and 0.650 for the shipped head. It serves only as H19's source. The shipped head is served through the retrained C″ arm, so this run says nothing about it.

H21 registered an abstention band meant to trade a small rise in handoffs for fewer wrong claims. The mechanically chosen band abstained on 1.8 percent of claims on ap-eval and 0.7 percent of claims on ap-challenge. Both figures sit below the run-to-run noise documented above for H19. In a live comparison against an independently rerun baseline, results moved by more than the band could have caused given how few claims it touched, and none of the cases whose claim status changed between the two runs were cases the band touched. H21 is reported as inconclusive. Its replacement is a counterfactual instrument that reads the band's effect from a single run's own trajectories, with no separately rerun baseline.

6.3 J5: the retrained head, fresh eval2 and the challenge split

Table 4. Decision-class accuracy, J5 (n = 800 fresh eval2 cases; n = 200 challenge cases).

split A (rules) C″ (retrained decision head) L (fine-tuned learned head) C (Jev, hosted)
eval2 (fresh, n = 800) 0.566 0.728 0.701 not run on eval2
challenge (n = 200) 0.565 0.739 0.688 0.536

C″ − A, eval2: +0.161, 95 percent CI [0.131, 0.191]. C″ − A, challenge: +0.171, 95 percent CI [0.115, 0.224]. L − A, challenge: +0.124, 95 percent CI [0.033, 0.211]. C″ − C (Jev), challenge, paired on the same cases: +0.199, 95 percent CI [0.143, 0.254]. L − C″, eval2: −0.027, 95 percent CI [−0.056, +0.002]. Every interval above excludes zero except this last one. That comparison fails this project's own pre-registered acceptance rule for the learned head, so the learned head ships as experimental and the retrained head stays the default.

6.4 Wrong-claim rate and the scorer defect

Table 5. Wrong-claim rate, challenge split (n = 200), registered exact-match scoring against a predicate-equivalent reading.

arm registered (exact document id) predicate-equivalent
A, rules 0.27 0.03
C, Jev 0.22 0.015
C′, zero-shot decision head 0.28 0.085
C″, retrained decision head 0.347 0.095
L, learned head 0.33 0.09

Error analysis traced 400 of the 435 wrong claims counted across every arm and split to one scorer defect, and an independent re-derivation from the raw case files confirmed it. The challenge generator attaches a second, independently valid document to a requirement that already has one on file, and the registered scorer's exact-id check marks the pack wrong for citing the other one. The predicate-equivalent column credits either valid document and scores every other mismatch, such as a stale document or a missing citation, wrong exactly as before. The defect is a measurement artifact in the harness and says nothing about how carefully the heads reason compared with the rule table. No decision-class accuracy number in Section 6.1 or 6.3 changes.

Table 6. Wrong-claim rate, fresh eval2 split (n = 800), before and after a pack-builder fix.

arm before, registered exact-id after, validated pack
A, rules 0.0112 0.0088
C″, retrained decision head 0.0466 0.0113
L, learned head 0.0537 0.0163
Wrong-claim rate before and after the evidence-pack fix
Figure 4. Wrong-claim rate before and after the evidence-pack fix. Share of cases closed with a claim the registered scorer marks wrong.

The 0.0466 reading sits above 0.04, the rate this project pre-registered as the trigger for redesigning the claim guard. A taxonomy of the judging head's 37 eval2 wrong claims, independently verified, found the cause. Thirty-one were packs the environment's own pack builder had padded with extra supporting documents beyond the ones the case key names; 91 of the 101 extra documents fail the harness's own validity predicate for entity and currency, the same check the ledger already applies elsewhere. Four were the worker citing the wrong document on an indirect requirement. The rule table makes that error on the same cases, so it is shared with the head. The remaining two were further citation mismatches. The taxonomy's independent verification traced none of the 37 to the head's own judgement question or to the elimination walk closing an item early.

The remedy was registered before it was computed and independently verified afterward. The pack builder now checks every supporting document for entity and currency before listing it, using the predicate the ledger already applied to decide whether an item was met. Under the registered exact-id scoring rule this brings the eval2 wrong-claim rate to 0.0088 for the rules, 0.0113 for the judging head, and 0.0163 for the learned head, down from 0.0112, 0.0466, and 0.0537. Every case that changed moved from wrong to verified, and no other case moved. Both heads meet the pre-registered threshold of 0.02 on the validated pack and miss it on the original pack, and this paper reports both readings side by side. We found the fix by tracing why the threshold was first missed. It is reported as a correction to the environment, made and disclosed after the fact.

6.5 Retiring the claim guard

Three approaches to a claim guard were tried after J5, and none earned a place in the design. A vote-based guard sampled several readings of the completion pack and voted. It separated wrong from verified claims at an AUC of 0.56, below its pre-registered gate. A deterministic evidence guard initially reported a 79-of-79 catch rate on the same kind of case. An adversarial review found that its entity and currency checks were the same two functions the case generator uses to define a wrong claim. A check scored against the rule it reuses cannot miss, so the number was withdrawn as circular. A reading-layer pilot then opened the cited document and compared it against the requirement, where the earlier guards had voted or checked dates. Combined with the deterministic layer, it produced numbers identical to the deterministic layer running alone, and on its own it separated a wrong claim from a right one at an AUC of 0.52. No model-based guard added anything past the deterministic check. All three guards are retired, and the deterministic validity check now runs inside Section 6.4's pack-builder fix, which validates every supporting document before a claim leaves. The check sits in the workflow engine's own bookkeeping and calls no model. Once it runs, the judging head claims wrongly at close to the rule table's own rate.

6.6 J6: transfer to a second process vocabulary

J1 and J5 measured one synthetic process, supplier onboarding. J6 asked whether the rule table, the judging head, and the hosted Jev service carry to a second process built with the same generator family: accounts-payable exception routing, on two fresh splits, ap-eval at 400 cases and ap-challenge at 200 cases. The AP generator shares onboarding's blocker mix and document counts by construction. Before any AP data existed, adversarial review relabelled H15 and H16 as claims about vocabulary transfer, meaning a new process vocabulary layered on the same structure. Transfer to a process with a different blocker mix and obligation structure is registered as future work, J10.

Table 7. J6 decision-class accuracy by split and arm, with the registered comparisons (case-clustered paired bootstrap, 10,000 resamples, seed 0).

split A C″ L-onb L-ap C (Jev) C″−A L-onb−C″ L-ap−C″ L-ap−L-onb C−C″
ap-eval (400) 0.648 0.650 0.522 0.663 0.636 +0.002 [−0.029,+0.034] −0.128 [−0.174,−0.081] +0.013 [−0.016,+0.043] +0.141 [+0.088,+0.192] −0.015 [−0.052,+0.024]
ap-challenge (200) 0.638 0.659 0.558 0.671 0.622 +0.022 [−0.023,+0.066] −0.101 [−0.158,−0.041] +0.011 [−0.034,+0.058] +0.113 [+0.038,+0.183] −0.037 [−0.088,+0.014]
The head's gap to the rule table, run by run
Figure 5. The head's gap to the rule table, run by run. Decision head minus rule table, decision-class accuracy. The first run used the zero-shot head.

H15, rejected. C″ beat A by +0.002 on ap-eval, far short of the registered 0.03 bar, and the interval straddles zero; on ap-challenge the same comparison is +0.022, also short of the bar. The judging head landed inside its predicted band, 0.650 and 0.659 against 0.58 to 0.68. The rule table scored 0.648 and 0.638, above its own predicted band of 0.52 to 0.60. The predicted gap closed because A rose while C″ held, which leaves the rules level with the head on this vocabulary. J7 is registered to test whether the AP generator produces cases whose evidence the rule table can already see in full, leaving the judgement question nothing to add.

H16, split. L-onb, the onboarding-trained head applied to AP with no retraining, came in at −0.128 against C″ on ap-eval and −0.101 on ap-challenge against a 0.05 band, with both intervals entirely on the wrong side of it. Its training on one process did not carry to the other. L-ap, retrained on 100 labelled ap-dev cases, recovered parity with C″ at +0.013 and +0.011, both inside the 0.03 band, and came in +0.141 and +0.113 above L-onb on the two splits, with both intervals clear of zero. On this evidence, a hundred of a company's own labelled cases (271 decision checkpoints here) brought the learned head level with the judging head on this vocabulary.

H17, reference. Jev scored 0.636 on ap-eval and 0.622 on ap-challenge, at or below both the rule table and the judging head; only the ap-challenge gap against L-ap clears zero. Table 8 compares claim and handoff rates on packs rebuilt through the same validity filter on both sides.

Table 8. J6 rebuilt-pack claim-rate and unnecessary-handoff comparison, Jev against the rules and the judging head, validated packs on both sides (case-clustered paired bootstrap, 10,000 resamples, seed 0).

split comparison wrong-claim (predicate) diff [95% CI] unnecessary handoff diff [95% CI]
ap-eval C − A −0.023 [−0.040, −0.005] +0.123 [+0.088, +0.158]
ap-eval C − C″ −0.018 [−0.035, −0.003] +0.138 [+0.105, +0.173]
ap-challenge C − A −0.055 [−0.095, −0.020] +0.125 [+0.080, +0.175]
ap-challenge C − C″ −0.030 [−0.060, −0.005] +0.130 [+0.085, +0.180]
Jev on the second process
Figure 6. Jev on the second process. Share of cases, Jev minus the judging head, with both sides scored on validated evidence packs. Left of zero means Jev does it less often.

Jev claims a case done wrongly 1.8 to 5.5 points less often than the rules or the head, and hands a case off unnecessarily 12.3 to 13.8 points more often, the same trade seen on onboarding in Section 6.4. Its accuracy on this process sits at or below the other two arms, so the lower wrong-claim rate comes from caution.

One caveat J6 raised is now closed. L-ap's training features and serving path both called C″'s own live judgement output, so its parity with C″ could have come from a classifier stacked on its own input. PREREG-J7 registered arm L0 to test this. It is the same learned head, trained on the same 100 ap-dev cases with every prior-answer feature removed, served with no call to the judging head. H22 ran on both AP splits and is independently verified in results/verification-J7-H22-2026-09-30.md.

Table 7a. J7 H22, the learned head with no judging-head input, decision-class accuracy against the stacked head and the reference arms (case-clustered paired bootstrap, 10,000 resamples, seed 0).

split L0 L0 − L-ap L0 − L-onb L0 − A L0 − C″
ap-eval (400) 0.660 −0.003 [−0.023, +0.016] +0.138 [+0.083, +0.189] +0.012 [−0.011, +0.034] +0.010 [−0.019, +0.040]
ap-challenge (200) 0.670 −0.001 [−0.031, +0.031] +0.112 [+0.037, +0.181] +0.032 [+0.005, +0.061] +0.011 [−0.033, +0.055]

L0 matches L-ap with no judging-head input, and the L0 − L-ap intervals straddle zero on both splits. On ap-eval L0 also claims wrongly less often than every reference arm, by 0.02 to 0.035 predicate points, with intervals clear of zero. A hundred labelled cases with structured and history features alone reach the accuracy of the stacked version, and the J5c and J6 circularity caveat on arm L is withdrawn. Read beside J8, the stacked classifier and L0 hold at 0.66 to 0.67 on the ap-dev split while an encoder fine-tune falls to 0.55. The encoder export carries the 201 checkpoints with open items from 58 of those cases; L-ap and L0 saw the same 201 plus the 70 terminal checkpoints that label DONE. On those same open decisions, feature design decided the comparison between the feature-based heads and the encoder.

A second caveat still limits what J6 licenses. The four onboarding-harness arms, A, C″, L-onb, and L-ap, ran on the registered pack builder without the addendum-8 validity filter, while arm C ran through the isolated harness with that filter already applied. Table 8 rebuilds all four arms' packs through the same filter before comparing, and only that rebuild makes the claim-rate and handoff numbers readable. Decision-class accuracy stays as it was, since the rebuild replaces a claim's supporting-document list and leaves the chosen action alone.

6.7 J8: two open alternatives on the same benchmark

Laya and CLM-8B, the two open decision models described in Section 2, were run on this project's own pre-registered benchmark, so that their measurement no longer rests on vendor claims alone. PREREG-J8 registered six hypotheses, H23 through H28, before any run. LY-ft is Laya fine-tuned on the training run the learned head used, exported as one row per open checkpoint: 1,567 six-way decisions and 2,901 per-item questions from 363 of the 600 onboarding training cases, and separately 201 decisions and 376 per-item questions from 58 of the 100 ap-dev cases. The cases the export leaves out were complete at their first checkpoint and carry no decision to learn from (j8_export_training.py; PREREG-J8 addendum 18). CLM-ft is CLM-8B with its two contrastive heads trained the same way. LY-0 and CLM-0 are both models as released, with no fine-tuning. Reference rows for A, C″, L, and Jev were already banked. Every arm answered the same rendered state the judging head answers. The case key, the generator's hidden blocker field, the oracle's expected action, and every judging-head prior answer were stripped before an encoder ever saw the text, checked by a selftest that greps the built training and serving text for those fields and asserts zero occurrences. Four splits ran: the two-hundred-case onboarding challenge split, the two-hundred-case ap-challenge split, the four-hundred-case ap-eval split, and a four-hundred-case seed-17 subsample of eval2. Each cell was independently recomputed from raw run-directory files within a tolerance of 0.0005, and each passed leakage, training-overlap and oracle-label audits before any hypothesis was read.

Table 9. J8 decision-class accuracy, four splits, verified (addenda 13–16).

split arm accuracy − A [95% CI] − C″ [95% CI]
challenge (200) LY-ft 0.815 +0.251 [+0.191, +0.307] +0.076 [+0.042, +0.112]
challenge (200) LY-ft seed 1 0.813 vs seed 0 +0.002 [−0.029, +0.043]
challenge (200) LY-0 (untuned) 0.583 +0.018 [−0.002, +0.039]
challenge (200) CLM-ft 0.687 +0.123 [+0.070, +0.176] −0.048 [−0.107, +0.010]
challenge (200) CLM-ft seed 1 0.563 vs seed 0 −0.125 [−0.217, −0.029]
challenge (200) CLM-0 (untuned) 0.537 −0.028 [−0.105, +0.051]
ap-challenge (200) LY-ft 0.554 −0.084 [−0.156, −0.009] −0.105 [−0.173, −0.037]
ap-challenge (200) LY-ft seed 1 0.560 vs seed 0 +0.006 [−0.062, +0.075]
ap-challenge (200) LY-0 (untuned) 0.653 +0.015 [−0.016, +0.046] −0.006 [−0.053, +0.039]
ap-challenge (200) CLM-ft 0.568 −0.070 [−0.140, +0.002] −0.091 [−0.151, −0.031]
ap-challenge (200) CLM-ft seed 1 0.570 stable
ap-challenge (200) CLM-0 (untuned) 0.656 +0.018 [−0.012, +0.047]
ap-eval (400) LY-ft 0.539 −0.110 −0.112 [−0.166, −0.057]
ap-eval (400) CLM-ft 0.586 −0.062 [−0.107, −0.016] −0.064
eval2 subsample (400) LY-ft 0.824 +0.276 [+0.232, +0.319] +0.069 [+0.049, +0.090]
eval2 subsample (400) CLM-ft 0.641 +0.093 [+0.051, +0.137] −0.111 [−0.158, −0.065]
Accuracy against the rule table on the harder set of each process
Figure 7. Accuracy against the rule table on the harder set of each process. Each dot is one approach's difference in decision-class accuracy from the rule table on the same cases. The right-hand column gives the accuracy itself. Fine-tuning and retraining used the 600-case onboarding training split and the 100-case accounts-payable development split.

References. Challenge: A 0.565, C″ 0.739, L 0.688, Jev 0.536. Ap-challenge: A 0.638, C″ 0.659, L-onb 0.558, L-ap 0.671, Jev 0.622. Ap-eval: A 0.648, C″ 0.650, L-onb 0.522, L-ap 0.663, Jev 0.636. Eval2 subsample: A 0.549, C″ 0.756 on 394 of 400 cases, L 0.735.

H23 predicted LY-ft within 0.05 of C″ on eval2 and below the rule table on challenge. The prediction fails on three of the four splits, in the encoder's favour on two of them. LY-ft beats C″ by 0.076 on challenge and by 0.069 on eval2. On ap-eval it trails C″ by 0.11, far from the predicted near parity. The prediction's direction holds only on ap-challenge, the split whose fine-tune saw the least data, where LY-ft falls below the rule table and inside its own band. H24 predicted both untuned encoders near the majority-class baseline, and both land level with the rule table on the two splits they ran. H25 predicted CLM-ft below the rule table on every split. It beats the rule table on challenge and eval2, points the predicted way without significance on ap-challenge, and is supported only on ap-eval. H26 predicted both fine-tunes would claim a case complete wrongly more often than C″, and wherever the comparison is significant the reverse holds. Both fine-tunes claim wrongly less often than C″ on eval2 and ap-eval. On challenge LY-ft ties C″ on wrong claims, while CLM-ft shows a significant reduction. Neither fine-tune differs from C″ on ap-challenge. H27, descriptive only, holds. Laya answers in 0.17 seconds. CLM-8B averages 0.4 to 0.6 seconds on the onboarding challenge split and 0.6 to 0.8 seconds on the AP splits, with a tail reaching about nine seconds on both. The judging head's fast mode takes a median of 3.1 seconds per decision. H28, calibration, is confirmed, as reported below.

Every reversal comes from the same class. On both splits where LY-ft beats C″, the whole advantage sits in REASON, with recall of 0.917 against C″'s 0.740 and the rule table's 0.104 on challenge, and 0.966 against 0.809 and 0.060 on eval2. With 1,567 labelled decisions, a 421-million-parameter encoder learns to recognise when a case needs judgement more accurately than the question-asking head infers it from the same rendered state. On the AP splits the same recipe, trained on only 201 labelled decisions, over-promotes INFO decisions to REASON at a rate of 0.53 for Laya and 0.39 for CLM-8B, against 0.02 and 0.06 untuned. INFO recall collapses and accuracy falls below the rule table it was meant to beat. Across the encoder's own two training sets, the 1,567-decision fine-tune beat the 201-decision one, so training size decided that comparison. The untuned encoders sit level with the rules on every split regardless of training size.

Where the gains and losses sit, by decision class
Figure 8. Where the gains and losses sit, by decision class. Share of checkpoints of each true class that an approach classified correctly. The DONE class is left out because the harness decides it without a model call and every approach scores 1.000.

Two calibration readouts close this section. H20, PREREG-J7 addendum 9, found the judging head's own item-level probabilities already calibrated on eval2, with expected calibration error of 0.0178 before any temperature refit and 0.0143 after one fitted on the training split. The 0.52 figure that started this line of work was a decision-class calibration reading, and that decision-class number stays open. H28 asked whether LY-ft's item-level calibration, after Laya's own vendor refit, beats C″'s uncalibrated reading on the same cases and comes within 0.05 of C″'s own refit reading. On the same four hundred eval2 cases, LY-ft's item-level ECE is 0.0057 uncalibrated and 0.0054 after its refit, against C″'s 0.0313 uncalibrated on the identical cases. Both registered conditions hold, and H28 is confirmed. The head was already well calibrated at the item level before this campaign ran, so Laya's margin over it is small in absolute terms.

What training on a process's own decisions did
Figure 9. What training on a process's own decisions did. Decision-class accuracy on the harder set of each process, 200 cases each.

Table 10. J8 scorecard, all splits verified (PREREG-J8 addendum 17).

hypothesis onboarding challenge eval2 (400) ap-challenge ap-eval
H23 LY-ft vs C″ wrong, encoder +0.076 wrong, encoder +0.069 supported (below A) wrong, trails by 0.11
H24 LY-0 near random wrong (level with A) not run wrong (level with A) not run
H25 CLM-ft below A wrong (+0.123) wrong (+0.093) direction right, not significant supported
H26 wrong claims above C″ wrong (tie) reversed (fewer) wrong (no difference) reversed (fewer)
H27 latency Laya 0.17 s, CLM 0.4–0.6 s onboarding / 0.6–0.8 s AP, heavy-tailed to 9 s, reported
H28 calibration confirmed

Four limitations, registered in PREREG-J8 addendum 17, qualify this section; Section 7 reports them together with every other campaign's.

6.8 What was not measured

J3, the expert-consultation phase and its blind-adjudication sample, has not run. This project's registered calibration function cannot score the learned head, because the head outputs a classifier's class probabilities and the function reads a per-item probability map. A design-time reading exists, but it was fitted on the same data used to build the model, so this paper leaves it out. Both gaps are reported as not yet measured, which carries no negative result.

7. Limitations

Every case in this study is synthetic, generated deterministically from templates with no model involved. That gives ground truth correct by construction and guarantees that no client evidence appears anywhere in this project. The cost is that we do not yet know whether the same distinction holds on the messier documents a real supplier file contains. Eval2 came from the same case templates and process family as the split the retrained head was tuned against. Its gains (Section 6.3) are a fresh sample from the same regime, and nothing here says whether they survive a different workflow. J6's transfer test carries the same limit to a second process. The AP splits share onboarding's blocker mix and document counts by construction, so the transfer results in Section 6.6 concern a new process vocabulary layered on the same structure. Whether these results hold on a process with a different blocker mix and obligation structure is registered as future work, J10.

The retrained head's wrong-claim rate cleared its pre-registered bar only after a defect in the environment's own pack builder was found and fixed. On the original pack it remains above the trigger. Because the fix came after the trigger was missed, both readings are reported side by side (Section 6.4). The learned head ships as experimental by the same pre-registered rule, having failed its own acceptance test against the retrained head it was meant to match. No model-based claim guard passed a fair test. The vote guard and the reading-layer pilot each added nothing beyond a deterministic check that reproduces the benchmark's own validity rule, the same check the pack-builder fix relies on (Section 6.5).

The worker and the decision head in arm C′ are both built on DeepSeek V4.1 Flash, served through a three-bit-per-weight quantization spanning two GPUs as one model. Another quantization, another base model, or an unquantized reference may calibrate and discriminate differently, and this registration says nothing about where on that curve the effect appears or disappears. The one-off probe and the service's own validation smoke, both against qwen3.8:27b (Section 3), showed only that a logprobs read is possible on a model of that kind, and J1 does not measure that model.

The decision head reads a single generated token per question, a narrow window onto what a model believes. A question whose answer needs more than one token cannot be asked this way, and this design makes no attempt to widen the window. The calibration measurement has a gap that was registered in advance. A correctly timed completion closes every obligation and produces no corresponding question, so H3 measures calibration on open and blocked cases and leaves out the moment a case correctly closes. Adversarial review also made a related structural fact explicit. In this version of the question set, CONTINUE and CONSULT_EXPERT are decided from one shared probability, so a controller cannot yet state its confidence that a case needs a stronger model separately from its confidence that it can handle the case itself. A question set that separates them is future work.

The counterfactual replay behind Section 6.2's H5 result carries a structural bias. replay() completes every branch from a checkpoint, including the one chosen, with the deterministic rules controller regardless of which arm is under test. Arm A benefits from this every time, because its continuation and the yardstick are the same policy. The H5 refutation holds as arithmetic under this instrument and ranks decision quality poorly. J7 registers a replay that completes each branch with the policy of the arm under test.

Section 6.2's consistency figure depends on the serving configuration as well as on the model. J7's H19 traced the open head's apparent inconsistency against Jev to the serving stack, first to ambient load from a second campaign sharing the endpoint and then to batching once concurrency rose. At one call in flight, on an endpoint with no other job attached, the open head disagrees with itself on 0.05 percent of repeats, matching Jev's own J1 figure exactly. The shipped service now defaults to one call in flight, which removes the gap at the cost of throughput; this project did not measure how far concurrency can rise before the figure moves, or whether Deliberate mode, Section 4's higher-cost setting that samples several answers per decision instead of one, closes the gap faster than serialising calls does. H21's abstention band, registered as a partial hedge against the disagreement that remains, abstained too rarely to be read from a live rerun against a separate baseline; whether it helps at all is still open, and a counterfactual reading of a single run's own trajectories is registered in its place.

A typed-question head should be kept away from arithmetic and date logic. The vendor's own documented failure modes for Jev name both, and this experiment's environment applies the same lesson by keeping currency and entity checks in a shared rule module that no model, hosted or local, is ever asked to reproduce.

J8's run against the two open encoders carries limits of its own, registered in PREREG-J8 addendum 17. Every case in J8 comes from the one synthetic generator family this project has used throughout, and the training labels and the scoring function share one oracle, the same limit the learned head carried in Section 6.6. The DONE class is a zero-call shortcut credited to every arm alike, encoder and rule table. The encoder recipes were reconstructed from the vendors' own published primitives without running vendor tooling end to end, and CLM-8B's encoder runs behind a hand-written transformers shim whose numerical parity with the vendor's own vLLM serving path is unverified. The runs hit three Metal-backend crashes on the serving side before a clean run. That was a serving-stack problem with no bearing on the modelling, of the same kind J7's H19 traced on the fleet's own endpoint. A later diagnostic, registered in PREREG-J9 and reproduced independently, bounds what the encoder learned. A four-feature rule computed from the rendered observation alone, with no learning and no document reading, drives the same elimination walk to 0.754 on the eval2 200-case subsample and 0.793 on the challenge split, between 0.02 and 0.045 below the fine-tuned encoder depending on the split. Part of the encoder's gain is a cue the state already carries, and J9 is registered to measure how large that part is. A second late diagnostic examines the shared oracle itself (PREREG-ORACLE, 6 October). A frontier model that had seen neither arm's answers judged 197 blind checkpoints from the case files alone. Where the oracle's rule and the file give one answer, it agreed every time, 28 of 28 INFO and 83 of 83 REASON. On the 120 checkpoints where the rule table and the head disagree, it sided with the oracle on 99, and the oracle's class there was the head's in 85 cases and the rule table's in 8. It never chose BLOCKED, because the fact that no counterparty can supply a document lives in the hidden key and is absent from the file. It also read every closed item that rested on a related-entity document as still needing interpretation, a difference in process definition with no model fault behind it. The same strict reading makes it agree with the registered scorer that 36 of 37 eval2 wrong claims were wrong. On the facts it agrees with Section 6.4's taxonomy, which found most of the extra documents in the padded packs failed the validity predicate; the two differ only on whether a related-entity document closes an item by interpretation, which the process rules allow and the adjudicator read strictly. Three checkpoints where the oracle itself may be wrong are flagged for a human look. In J8 training size decided the encoder-against-encoder comparison and feature design decided the head-against-encoder comparison, and a learned mode must be checked on held-out data before it is promoted. The judging head remains the default. It reads the cited evidence, which neither encoder does, and it ships with a Jev-shaped API and an audit row behind every decision.

8. Reproducibility

Every module in this project (case generator, environment, controllers, scorer, runner, and the decision head itself) is written in Python 3.11 against the standard library only, so the harness runs anywhere a plain Python install does. Ground truth is correct by construction. The case generator builds the hidden key first and then renders the documents, emails, and application data from that key, so no data is labelled after the fact. The frozen question set and decision table carry a recorded SHA-256 hash checked at run time (questions_v1.json, version 1.2), and the harness refuses any run against a different question set. Every adversarial review is pinned to a specific commit, and every module carries its own self-test, printed as an N-of-N pass count, that must pass before its output is trusted downstream. The worker model is pinned per run in runs/<run>/config.json; a change of worker, such as the move to DeepSeek V4.1 Flash, starts a new run so that no run mixes two workers.

Every phase reported in Section 6 was independently verified from raw run files by a party who did not build the scorer, in each case importing the scorer's own metric functions without re-implementing their logic. J6's transfer arms, Jev's read-out, and the rebuilt-pack rescoring behind Table 8 are each verified in their own file (results/verification-J6-2026-09-26.md, results/verification-J6-armC-2026-09-27.md, results/j6-rebuilt-packs-2026-09-27.md). Section 6.2's H4 and H5 numbers were verified twice, the second time by direct recomputation from the raw repeat and branch files (results/verification-J1-H4H5-raw-2026-09-28.md); every cell matches the numbers reported in Section 6.2 within the stated tolerance. H19 and H21 were verified the same way, with all 28 of 28 cells matching the launcher's own numbers within tolerance 0.0005 (results/verification-J7-H19-H21-2026-09-29.md). H22 was verified by recomputing every cell from raw trajectory files and checking the model's own provenance fields, hash, feature list, and fast-head call count (results/verification-J7-H22-2026-09-30.md). J8's four splits are each independently verified in their own file (results/verification-J8-challenge-2026-09-29.md, results/verification-J8-ap-challenge-2026-09-29.md, results/verification-J8-ap-eval-2026-09-29.md, results/verification-J8-eval2-2026-09-30.md), each within tolerance 0.0005 and each additionally audited for leakage, training-set overlap, and the oracle-label limitation before a hypothesis was read. Section 6.7's two calibration readouts are each in their own file (results/j7-h20-calibration-2026-09-30.md, results/j8-h28-calibration-2026-09-30.md).

Ethics and data

No client data of any kind appears anywhere in this project. Every case, document, and email is generated locally and deterministically from a synthetic policy and synthetic supplier templates. A hosted model is called in only two places, a narrowly scoped expert-consultation step and a blind adjudication sample. Both are funded under a fixed credits ceiling, and an explicit rule restricts them to judging and evaluation, so neither generates any of the material the system is tested on. Those two calls are the only traffic that leaves the local network, apart from calls to the hosted Jev service through OpenRouter, which are logged and reported as a separate cost line.

Acknowledgments

The author's automated research fleet built the harness, ran the three adversarial reviews, and implemented and self-tested the decision head, working from a brief and two internal working papers. The internal papers are cited here only as the origin of the question and are not public sources.