ContentsAll campaigns
Tinholt LabAnnual report · 2026

Pre-registered measurement of delegation and oversight

The Measured
Enterprise 2026

Organisations are handing decisions to AI agents faster than they are building the oversight to match.

The lab measures where delegation breaks and what oversight buys back, in pre-registered campaigns on local models, and publishes the result whichever way it falls.

90,880organisation designs swept for the price of delegation depth
30,000supervised episodes across five irreversibility tiers
18,360units audited across all fault classes, including 3,240 real diaries
100simulation runs of first-mover AI adoption across four industry verticals
0mismatches when 14,950 convex and 14,952 linear scored episodes, 29,902 in all, were recomputed independently

Dinand Tinholt

Head of AI Center of Excellence, Americas, Capgemini

DraftFirst full draft · 6 October 2026

The Measured Enterprise 2026

Contents

  1. SummaryExecutive summary3
  2. Chapter 1The Delegation Index5
  3. Chapter 2How the lab works9
  4. Chapter 3AGENESIS: first-mover adoption11
  5. Chapter 4Error Independence12
  6. Chapter 5Delegation Cliff14
  7. Chapter 6SIGIL16
  8. Chapter 7AGENESIS-2: sprawl cost18
  9. Chapter 8Trial Balance20
  10. Chapter 9The open decision head22
  11. Chapter 10What comes next26
  12. Chapter 11Corrections logged27
  13. Appendix AMeasured claims index29
  14. Appendix BGlossary32
  15. Appendix CSources34

Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Charts are rebuilt from those numbers.

06 OCT 2026

Summary

Executive summary

The reading

The Delegation Index reads 47.2 for the third quarter of 2026, with a band of 40.8 to 53.2. Its inputs are the lab's own synthetic measurements, the simulations and controlled model runs on generated tasks from five published campaigns, 139,646 units in all. It says nothing about the market or about any single firm. The number summarises how those campaigns came out when enterprise work went to AI agents under oversight and the oversight budget stayed fixed. At a reading of 50 the two balance under the conditions the lab measured.

47.2Delegation Index, Q3 2026. Band 40.8 to 53.2. Five campaigns, 139,646 units.

This first reading sits 2.8 points under balance, and its band contains 50. This quarter's evidence cannot tell balance from imbalance, which is why the number is always quoted with its band.

Five components

Each component reads one published number and scores it from 0 to 100. All five weigh the same.

ComponentCampaignPublished number it readsScoreBand
Delegation priceDelegation Cliffdepth −1.499 against reviewer capacity +1.45649.3none published
Oversight yieldSIGILvalue of plain review against no oversight, five tiers46.946.8 to 47.0
Audit detectionTrial Balancechecklist reader at a full read, 0.58358.347.9 to 66.7
Reviewer independenceError Independencemean cross-family error correlation, −0.10255.136.0 to 73.5
SprawlAGENESIS-2cost exponent under a flat monitoring budget, 1.466526.723.9 to 29.5

Sprawl holds the index down. With the monitoring budget flat, oversight cost grows as fleet size to the power 1.47, and that leg scores 26.7. Audit detection lifts it, because a reviewer who read the agent's whole record with a per-step checklist found 0.58 of the skipped steps. The other three land within five points of balance, and Chapter 1 walks through the arithmetic step by step.

What the campaigns found

Error Independence ran one model on two different machines and found that the copies share their mistakes, with an error correlation of +0.447 on the items where the models disagreed. Pairs from different model families averaged −0.102.

Delegation Cliff swept 90,880 simulated organisations. Each added level of delegation depth cost 1.50 on the quality margin, and reviewer capacity bought back the most at +1.46. The quality floor decided whether a cliff showed up at all. It took a floor of 0.99 to turn most of the space fragile; at 0.95 there was none.

SIGIL tested a prediction from the Accountability Rent working paper that review before commitment pays on irreversible decisions. The prediction failed in all 12 registered series. Give the reviewer a ledger of its own past damage and wrongful blocking at the top tier went from 30.1% to 42.7%. The lab published that as a refutation.

AGENESIS-2 measured what a fleet of agents costs as it grows. Under a flat monitoring budget the cost exponent is 1.4665, with a 95% interval of 1.4101 to 1.5230. Give every agent its own check and it falls to 0.9001. Before publication the independent check turned up duplicate runs, which moved one verdict from supported to refuted.

Trial Balance gave the auditor a ledger of what the agent owed. Reading only where that ledger failed to balance found 0.80 of skipped steps at 745 tokens. Log review topped out at 0.58, whatever the budget.

AGENESIS, published in April, simulated first-mover adoption and found a premium of 210%, with the market tipping in 77% of markets.

The product result

The year's product is the open decision head, a small service that answers typed questions about a live case with a probability for each answer, on a model the operator runs. Its first version failed to beat a hand-written rule table, with a gap whose interval included zero. A retrained version, tested on 800 fresh cases, scored 0.728 against the rule table's 0.566, a gap of +0.161 with a 95% interval of 0.131 to 0.191. At first it wrongly closed cases more often than the project's bar allowed. The cause sat in the evidence pack the environment built, and Chapter 9 prints both readings.

Corrections, and what comes next

Chapter 11 lists every correction the pieces report, among them the duplicate runs in AGENESIS-2 and a retracted sentence in Error Independence. Schouw is scheduled for 13 October. The next Delegation Index reading freezes on 5 January 2027.

05 OCT 2026

Chapter 1

The Delegation Index

What it measures

The Delegation Index is one number a quarter, on a scale from 0 to 100. It sums up what the lab's own published campaigns say about how far enterprise work can go to AI agents under oversight when the oversight spend stays fixed. At 50, the cost of handing work over matches what oversight buys back at no extra spend.

The construction was registered on 5 October 2026, before the Q3 numbers went through it. Only numbers printed in a published piece with a live public link on the freeze date enter. Anything internal stays out, drafts and unposted corrections included.

Version
1, registered 5 October 2026
Freeze
5 October 2026, Chicago time
Campaigns in
5 of 6 published
Recompute
independent, 6 October 2026, every value reproduced to two decimals

The selection rule

Each component reads the lab's best measured design in its campaign with oversight spend held fixed. The rule went on paper in advance so nobody could pick the flattering number afterward.

  • Delegation price expresses one increment of delegation depth in increments of reviewer capacity, the largest positive lever in the Delegation Cliff space. Both coefficients sit on the same band-standardised scale, so the ratio has no units.
  • Oversight yield uses plain review, the better of the two designs SIGIL measured before commitment, under the registered convex cost model.
  • For audit detection the rule picks the checklist reader at a full read in Trial Balance, the best diary audit that has a pooled rate and an interval in print.
  • Reviewer independence reads the five cross-family pairs in Error Independence, since that was the best pairing measured.
  • Sprawl takes the fixed monitoring budget in AGENESIS-2. No other measured arm keeps spend flat while the fleet grows.

The formula

Every component is scored from 0 to 100 and clipped to that range. Each balance point comes straight from the component's own definition.

CodeFormulaScores 50 when
P100 × capacity ÷ (capacity + depth)one increment of capacity buys back exactly one increment of depth
Y50 × (1 + m), m the mean over tiers of value ÷ stakereview breaks even
A100 × detection ratethe audit finds as many skipped steps as it misses
R50 × (1 − φ)the second reviewer's errors are uncorrelated with the first's
S50 × (2 − β)oversight cost grows in proportion to the fleet

The index is the plain mean of the five. Its band is the mean of the five lower ends and the mean of the five upper ends, which is the range the index would take if every component sat at the same end of its band at once. It has no coverage level, so the lab doesn't call it a confidence interval.

Figure 1 The five components and the index, Q3 2026. Sprawl holds the reading down; the index band contains 50.

0255075100balance, 50Delegation price, PDelegation price, P: 49.3 (no band)49.3 no bandOversight yield, YOversight yield, Y: 46.9 (46.8 to 47.0)46.9Audit detection, AAudit detection, A: 58.3 (47.9 to 66.7)58.3Reviewer independence, RReviewer independence, R: 55.1 (36.0 to 73.5)55.1Sprawl, SSprawl, S: 26.7 (23.9 to 29.5)26.7Delegation IndexDelegation Index: 47.2 (40.8 to 53.2)47.2
Component scores on the 0 to 100 scale with their bands. Bands come from a published interval (A, S) or the published spread of seeds or pairs (Y, R); P has none. The index band is the mean of the band ends and carries no coverage level.Source: The Delegation Index, Q3 2026 reading, registered, publication pending; inputs from the five published pieces cited in Chapter 1. Rebuilt from the published numbers.
The numbers as a table
ComponentScoreBand
Delegation price, P49.3no band
Oversight yield, Y46.946.8 to 47.0
Audit detection, A58.347.9 to 66.7
Reviewer independence, R55.136.0 to 73.5
Sprawl, S26.723.9 to 29.5
Delegation Index47.240.8 to 53.2

The arithmetic

Delegation price, P. "Pricing Agent Autonomy" prints a depth coefficient of −1.499 and a reviewer-capacity coefficient of +1.456, on the band-only Sobol fit, n = 4,948. P = 100 × 1.456 ÷ 2.955 = 49.27. One increment of depth costs 1.03 increments of reviewer capacity. The piece prints no interval, so P enters at its point value and adds no width to the band.

Oversight yield, Y. The SIGIL paper prints the value of review against no oversight per tier and per seed. The pooled values run −0.13, −0.17, −0.18, −0.17 and −2.64 from tier 1 to tier 5. Dividing each seed mean by its stake, the square of the tier, and averaging over the five tiers gives m = −0.062, so Y = 46.90. Plain review lost about 6% of the stake per decision. The three seeds give a range of 46.78 to 47.01.

Audit detection, A. Trial Balance prints omission detection at a full read with a per-step checklist of 0.583, 28 of 48, with a 90% cluster-bootstrap interval of 0.479 to 0.667. A = 58.3, with a band of 47.9 to 66.7.

Reviewer independence, R. Error Independence prints the error correlation of each pair on the 47 contested items. The five cross-family pairs read +0.280, +0.013, −0.105, −0.227 and −0.469, with a mean of −0.102. R = 50 × 1.1016 = 55.08. The spread of the five pairs gives a range of 36.00 to 73.45, the widest band in the index.

Sprawl, S. AGENESIS-2 prints an exponent of 1.4665 with a cluster-robust 95% interval of 1.4101 to 1.5230. S = 50 × (2 − 1.4665) = 26.68, with a band of 23.85 to 29.50. A higher exponent gives a lower score, so the band flips.

The index. (49.27 + 46.90 + 58.30 + 55.08 + 26.68) ÷ 5 = 47.25, which rounds to 47.2. The lower ends average 40.76 and the upper ends 53.19.

Behind the number

Five of the lab's six published campaigns enter. Their scale is 139,646 units: 90,880 configurations, 30,000 supervised episodes, 18,360 audits, 150 generated invoices and 256 simulation units. The units behind the five quoted numbers come to 20,154. Those units are different kinds of object, so the total is only a count. The two thinnest legs, 48 checklist audits and 47 contested invoices, carry the two widest bands.

Before Trial Balance entered, the other four components averaged 44.5, with a band of 39.0 to 49.8. Trial Balance went live on 1 October, inside the freeze window, and adding its audit component took the reading to 47.2. So all of that move comes from which campaigns are in.

Sensitivities

  • Oversight yield under the linear cost model scores 45.66, which moves the index to 47.0.
  • Reviewer independence read from the same-weights pair, +0.447, would score 27.7, 1.0 above sprawl at 26.7.
  • Audit detection read from the commitment-ledger audit at its lowest published cell, 0.92, would move the index to 54.0. That design has no pooled rate with an interval in print, so the registered rule leaves it out. It is printed here so readers can see it.
  • Sprawl under per-agent monitoring, exponent 0.9001, would score 55.0. That arm raises spend with every agent, which breaks the fixed-spend condition.

What stays out

AGENESIS measures competitive dynamics between firms and prints no oversight quantity, so it stays out. Trial Balance's commitment-ledger audit is the strongest design the lab has measured; once a published piece prints a pooled rate with an interval, it enters through a new version. In SIGIL's after-the-fact audit arm, the paper itself calls the crossings post-hoc and unstable across seeds. The Delegation Cliff quality-floor shares describe the shape of the design space, and the price component already reads that campaign.

What it does not claim

The index reads one lab's synthetic measurements under the lab's own conditions. The Delegation Cliff coefficients inside it are descriptive slopes. SIGIL and Trial Balance each ran one model in the measured seat, so the reading holds for those models. The balance point at 50 marks where the lab's measured conditions cancel. Outside this index it means nothing.

Rules that bind every reading

  1. Only numbers from live public pieces on the freeze date enter.
  2. The first paragraph of every release says the index reads the lab's own synthetic measurements.
  3. The headline never appears without its band and the count of campaigns and units behind it.
  4. Components with no band are named, and the release says the printed band is too narrow by their width.
  5. When the band contains 50, the release says the reading cannot tell balance from imbalance.
  6. The construction changes only through a new dated version, and the first release under it restates every earlier reading.
  7. Someone who did not build the reading recomputes it from the construction page and the cited pieces before release.
  8. A refuted result enters on the same terms as a supported one. SIGIL's refutation sits in oversight yield at full weight.
  9. The builder had read every input before fixing the balance points, because the inputs were public. The registration binds every reading after the first, which was made with the inputs already known.

06 OCT 2026

Chapter 2

How the lab works

The protocol

The lab studies how organisations hand work to AI agents and how they oversee it. Every campaign follows the same eight steps. If the lab tried to steer a result toward the answer it hoped for, the steps would leave a dated trace of the attempt.

  1. Ask one question.

    A campaign starts from a question a buyer or a regulator would recognise, such as whether a second model checks the first or what a liability ledger does to a reviewer. It names the quantity that would answer the question and the null result that would mean the idea adds nothing.

  2. Register the prediction, the bands and the kill conditions.

    Before any unit runs, the lab writes down the hypotheses and the band each number has to land in to count. The statistics and the conditions that stop the run go into the same document, which is dated and frozen. A later change goes in as a dated amendment, written before anyone reads the data it touches. A small calibration stage first checks that the task is neither trivial nor impossible. That gate checks the instrument works and nothing more. The lab learned where that line sits by halting a campaign whose gate had demanded the result it existed to measure.

  3. Attack the design.

    A reviewer whose job is to break the headline goes over the design and the analysis. Confounds and statistics that hide a split are the usual targets, along with ceilings and scoring rules that favour one arm. Whatever the review finds gets fixed before the run or stated as a limit in the piece.

  4. Run to the registered rule.

    The campaign runs on open models on local hardware the lab owns, so anyone with the code can repeat a run. It stops where the registration says it stops. A failed gate or a spent compute budget ends the run, and the lab analyses whatever is complete at that point. A run going badly keeps its length and its thresholds.

  5. Verify from the raw files.

    A separate pass with fresh code and no access to the original analysis recomputes the reported numbers from the raw records. Each piece says how many numbers were checked and how many matched. Where the two passes disagree, someone chases the gap down, and the piece carries the verified value.

  6. Publish the negative results.

    A refuted hypothesis appears with the same care as a confirmed one, next to the prediction it contradicted. A null result is reported as a bound. Each piece keeps what was measured apart from what the lab infers. An explanation stays labelled as one until a test has run against it.

  7. Release the code.

    Each campaign keeps its registration, harness, unmodified ledger and analysis together, so a reader can trace every published number to the rows behind it. The code and records go out with the pieces as each campaign is cleared for release.

  8. Log every correction.

    When a published number turns out wrong, the correction goes up in the open with a date, and the original stays visible beside it. The measured claims index in Appendix A keeps a row for every headline number, so a correction has a fixed place to land.

What the steps caught this year

During 2026 the verification pass on AGENESIS-2 found 95 duplicate pairs of simulations and moved a verdict from supported to refuted before the paper went out. Trial Balance's first attempt halted at its calibration gate with zero production units, and the gate rule was rewritten the same evening. SIGIL's registered prediction failed in every series, and it went out as a refutation. Chapter 11 lists every correction the pieces report.

The house format

The campaign chapters that follow share one layout. The question comes first, then what was registered before any unit ran, then what was measured and at what scale. The finding carries its number and, where the piece prints one, its interval. A chart rebuilt from the published numbers follows. Each chapter closes with what the campaign does not claim and with where the numbers were checked.

04 APR 2026

Chapter 3

AGENESIS: first-mover adoption

Published
4 April 2026, Medium
Scale
100 simulation runs, 4 industry verticals
Verdict
measured
Pre-registration
not recorded

The question

What happens to an industry when one company adopts AI agents before its competitors do? AGENESIS was the lab's first published campaign, and it asked that question of a simulated market.

What was registered

The site records no pre-registration for AGENESIS and no verification status. The findings are in the piece. This chapter reports them with no registration line, since none exists on the record.

What was measured

AGENESIS is an agent-based simulation of enterprise AI adoption. The published article reports 100 independent simulation runs across 4 industry verticals, under 5 adoption configurations and 3 macroeconomic conditions.

The finding

210%First-mover premium across 100 runs. The market tipped in 77% of markets.

The firm that adopted first earned a premium of 210% over its competitors, and the market tipped toward one firm in 77% of the simulated markets. The lab's later campaign on sprawl cost reused the same simulator and added an agent lifecycle to it.

What it does not claim

AGENESIS has no agent lifecycle. Agents are created once and never retired, and the AI cost a firm pays has nothing to do with how many agents it runs. A firm with fifteen agents and a firm with a hundred and fifty pay the same at the same level of adoption. The sprawl-cost campaign in Chapter 7 was built to fill that gap. AGENESIS prints no interval and no oversight quantity, so the Delegation Index leaves it out.

Where it was verified

The site records no verification status for this piece, and the two numbers above are quoted from the published article.

02 SEP 2026

Chapter 4

Error Independence

Published
2 September 2026, Medium
Scale
150 items, four model configurations
Verdict
measured
Ground truth
generated with each item

The question

A common resilience design runs the same model a second time, on different accelerators or in a second region, and treats the second answer as a check on the first. Civil aviation learned long ago that two identical computers fail the same way at the same moment. Flight control systems run on processors from different makers, with software written by teams kept apart, and the industry calls that dissimilar redundancy. This study asked whether a second copy of a model, on different hardware, makes different mistakes.

What was registered

Ground truth is generated with each item, so the correct answer is known before any model sees it. No model grades another, and so there is no judge inside the measurement. The site records no pre-registration beyond that design. For each pair of models the study computed the phi coefficient on error indicators. Phi sits near zero when two models go wrong on different items and near one when they go wrong together.

What was measured

Four model configurations scored the same 150 invoice-extraction items. The key comparison ran model A twice, in two builds with identical weights. The builds differed in quantisation and in the silicon and serving stack beneath them. Models B and C came from other families.

103 of the 150 items were unanimous. On 54 of them all four models were right. All four missed the other 49. Those items measure how hard an item is, so the headline uses the 47 items where the models disagreed.

The finding

Before conditioning, every pair looked correlated, from 0.556 to 0.827. On that reading you would give up on a second reviewer altogether.

+0.447Error correlation, identical weights on different silicon, 47 contested items. Cross-family pairs averaged −0.102.

On the contested items, the pair with identical weights stayed at +0.447, the highest in the matrix. The five pairs from different model families averaged −0.102, and model A's first build against model B sat at −0.469. Changes of machine and serving setup left the correlation standing. Only a change of model family moved it through zero.

Figure 2 Pairwise error correlation before and after conditioning on item difficulty. Only the pair with identical weights stays high.

−1.0−0.50.00.51.0open ring: all 150 items · dot: 47 contested itemsSame weights, two machinesSame weights, two machines, all 150 items: 0.827Same weights, two machines, 47 contested items: 0.4470.4470.827Model A build 2 + model CModel A build 2 + model C, all 150 items: 0.810Model A build 2 + model C, 47 contested items: 0.2800.2800.810Model A build 1 + model CModel A build 1 + model C, all 150 items: 0.729Model A build 1 + model C, 47 contested items: 0.0130.0130.729Model B + model CModel B + model C, all 150 items: 0.617Model B + model C, 47 contested items: −0.105−0.1050.617Model A build 2 + model BModel A build 2 + model B, all 150 items: 0.607Model A build 2 + model B, 47 contested items: −0.227−0.2270.607Model A build 1 + model BModel A build 1 + model B, all 150 items: 0.556Model A build 1 + model B, 47 contested items: −0.469−0.4690.556
Pairwise error correlation (phi) on all 150 items (open ring) and on the 47 contested items (filled dot). The highlighted pair runs identical weights on different hardware.Source: “When the Second Opinion Shares the Blind Spot”, Medium, 2 September 2026. Rebuilt from the published numbers.
The numbers as a table
all 150 items47 contested items
Same weights, two machines0.8270.447
Model A build 2 + model C0.8100.280
Model A build 1 + model C0.7290.013
Model B + model C0.617−0.105
Model A build 2 + model B0.607−0.227
Model A build 1 + model B0.556−0.469

What it does not claim

  • The same-family reading comes from one pair, the only two systems in the set with identical weights. One pair counts as an observation, and it would take more pairs before anyone quotes it as a rate.
  • The five cross-family pairs run from −0.469 to +0.280. A mean of −0.102 sits inside a spread that wide and describes it poorly. Two families can still fail together; the reading covers these families on this task.
  • The study ran one task, chosen because its ground truth cannot be argued with. Open-ended work is untested.
  • The article also reported a model-free checker as a floor for independence. A later note retracted that sentence. Chapter 11 gives the details.

Where it was verified

Model scoring ran through a separate verification path, and every model-pair correlation survived the later review that retracted the checker sentence.

03 SEP 2026

Chapter 5

Delegation Cliff

Published
3 September 2026, Medium, as "Pricing Agent Autonomy"
Scale
90,880 configurations, zero error units
Verdict
mixed
Certificate
a 200-pair certificate passed, 95th percentile 0.0117

The question

Any organisation deploying AI agents is betting that pushing work further from human review costs less than keeping it close, and few of them put a price on that bet. This campaign tried to. Where does delegation stop paying, and which design choices buy the margin back?

What was registered

Four predictions went on the record before the sweep.

  • H1. The design space has a cliff, a sharp boundary between robust and fragile configurations.
  • H2. Better observability outranks raw reviewer capacity.
  • H3. Agent self-check is the dominant lever.
  • H4. An audit death spiral forms, where eroding trust raises audit load until capacity saturates and quality falls.

What was measured

The campaign simulated an agent organisation with twelve design dimensions, and this chapter turns on two of them, delegation depth and reviewer capacity. The sweep scored 90,880 configurations on the margin between the quality each organisation produces and the floor it must hold. It finished with zero error units. A revalidation certificate reran 200 random configurations from scratch: the median difference was 0.002 and the 95th percentile 0.0117, against a tolerance of 0.05.

The finding

−1.50Change in quality margin per added level of delegation depth. Reviewer capacity buys back +1.46.

Delegation depth carries the largest weight in the space, with a negative sign: each increment costs 1.50 on the margin. Reviewer capacity buys back the most at +1.46. Model capability at +0.95 and self-check calibration at +0.92 follow, close enough to read as tied. Task coupling is the second-largest cost, at −0.96.

Figure 3 Linear-probe coefficients on the delegation margin. Delegation depth costs the most; reviewer capacity buys back the most.

−2.0−1.5−1.0−0.50.00.51.01.5Delegation depthDelegation depth: −1.499−1.499Reviewer capacityReviewer capacity: 1.4561.456Task couplingTask coupling: −0.958−0.958Model capabilityModel capability: 0.9490.949Self-check calibrationSelf-check calibration: 0.9150.915Rework costRework cost: −0.803−0.803Verification depthVerification depth: 0.4130.413Verification coverageVerification coverage: 0.3080.308Detection lagDetection lag: −0.259−0.259Workload volatilityWorkload volatility: −0.147−0.147Trust responseTrust response: 0.1020.102Escalation latencyEscalation latency: −0.007−0.007
What moves the delegation margin: linear-probe coefficients on the Sobol sample, band-only, n = 4,948. Positive buys margin, negative spends it.Source: “Pricing Agent Autonomy”, Medium, 3 September 2026. Rebuilt from the published numbers.
The numbers as a table
Design dimensionCoefficient
Delegation depth−1.499
Reviewer capacity1.456
Task coupling−0.958
Model capability0.949
Self-check calibration0.915
Rework cost−0.803
Verification depth0.413
Verification coverage0.308
Detection lag−0.259
Workload volatility−0.147
Trust response0.102
Escalation latency−0.007

The quality floor decides the failure geometry. At a 0.95 floor, 60% of the space is a broad transition band with no cliff. Raise the floor to 0.99 and the cliff appears, with 79% of the space fragile.

refutedH1 at the 0.95 floor. H2, H3 and H4 not supported.

The observability prediction failed. Capacity sits at +1.46, verification coverage at +0.31 and detection lag at −0.26. Interaction terms between observability and capacity came out real and positive, which makes observability a complement to reviewer capacity, and the hypothesis fails all the same. As for the spiral, mean margin rises across trust-response quintiles, even inside the high-stress stratum where depth and coupling are both high, so it never formed.

What it does not claim

  • The coefficients are descriptive slopes on the Sobol sample inside the transition band, n = 4,948. They describe how the margin moves where a boundary exists, and they carry no causal weight over the whole space.
  • The 0.99 result rests on one 256-unit calibration sweep. The 0.95 result is replicated on the 8,192-unit Sobol set. The article gives the two legs different weight for that reason.
  • Reviewer capacity is a parameter in this design. Reviewers here never tire.
  • Capability and self-check differ by 0.034. The article reads them as tied.

Where it was verified

Every result file carries a finite margin and a null error code, which is how the zero-error count was checked. The 200-pair revalidation certificate passed. The published coefficients come from the corrected band-only fit; Chapter 11 describes the correction that produced them.

10 SEP 2026

Chapter 6

SIGIL

Published
10 September 2026, Medium, as "Accountability Makes Oversight Worse"
Scale
30,000 supervised episodes
Verdict
refuted
Recompute
0 mismatches in 29,902 scored units

The question

Once a reviewer holds the power to stop an AI worker, the usual next step is to make it accountable for what it lets through. This campaign asked whether an AI supervisor authorises differently when it carries a running record of the damage its own past approvals caused. It also asked how the value of its review changes as decisions get harder to undo.

What was registered

The hypothesis came from the Accountability Rent working paper. It predicted that review before commitment loses value on reversible decisions, turns positive past some level of irreversibility, and climbs steeply from there. A flat response falsifies it, as does a concave one or the absence of any crossing.

The output contract was frozen on 30 August 2026, before any harness code existed. It fixed the scoring and a gate that required the unsupervised worker to err between 10% and 35% on a pilot. It also made a rerun under a second, linear cost model mandatory. One amendment, recorded before the second pilot, changed only the task generator and harness capacity and pre-committed a stopping rule against retuning more than once.

What was measured

An AI worker approved or rejected synthetic firm decisions across five tiers, from invoice approvals that reverse next cycle to public statements that cannot be taken back. The same base model sat in a second seat and reviewed it under four designs:

  • no oversight;
  • an after-the-fact audit of a 20% sample;
  • review with the power to block before commitment;
  • the same blocking review carrying a liability ledger of the harm its own past approvals caused.

Worker and reviewer share one model on purpose, so any difference comes from the incentive design. Every decision was scored against programmatic ground truth. The campaign ran 30,000 episodes on local hardware, in two full campaigns, one per cost model, with three seeds per cell.

The finding

refutedThe registered prediction failed in all 12 series, under both cost models.

Review never crossed zero, in any of the 12 registered series. Its value collapsed at the most irreversible tier, where the theory wanted it to rise, to −2.64 per decision against no oversight. The liability ledger made it worse, at −3.81.

42.7%Tier-5 decisions wrongfully blocked with a liability ledger. Plain review blocked 30.1%.

At that tier plain review wrongly blocked 30.1% of decisions, and the ledger pushed the share to 42.7%, since refusing was the one action it never charged for.

Figure 4 Net value of blocking review against no oversight by irreversibility tier. Both designs stay below zero and fall at tier 5; the liability ledger falls further.

0.0−1.0−2.0−3.0−4.012345Irreversibility tierno oversightreview, tier 1: −0.13review, tier 2: −0.17review, tier 3: −0.18review, tier 4: −0.17review, tier 5: −2.64−2.64 plain reviewledger, tier 1: −0.14ledger, tier 2: −0.20ledger, tier 3: −0.16ledger, tier 4: −0.51ledger, tier 5: −3.81−3.81 liability ledger
Net value per decision against no oversight, by irreversibility tier, convex damage model, pooled over three seeds. Below zero, oversight costs more than it saves.Source: “Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers”, Medium, 10 September 2026. Rebuilt from the published numbers.
The numbers as a table
Irreversibility tierplain reviewliability ledger
1−0.13−0.14
2−0.17−0.20
3−0.18−0.16
4−0.17−0.51
5−2.64−3.81

A reviewer that can only block can fix one kind of mistake, a wrongful approval. The share of worker errors that were wrongful approvals fell from 93% at tier 1 to 1.4% at tier 5, one false approval against 68 false rejections.

What it does not claim

  • This is a controlled simulation with one open-weight model family in both seats, scored against constructed ground truth.
  • The liability ledger is a thin, simulated form of accountability.
  • A follow-up registered to test stronger models stopped before it ran an oversight trial. Four other models got between 97.3% and 100% of cases right unsupervised, which left a reviewer nothing to catch, so the result holds for this setting only.
  • The after-the-fact audit arm shows positive cells in some series. Its seeds disagree on where and whether it crosses zero, and the paper reports those cells so nobody has to dig for them.
  • The contract-clause tier produced no worker errors across both campaigns, so its value figures are pure oversight cost.

Where it was verified

A separate verifier recomputed every table from the raw episode ledgers and found zero mismatches on 14,950 convex and 14,952 linear scored units, 29,902 in all. It also confirmed that the grid was complete and that no review unit carried ledger state. That pass surfaced one documentation defect in the tail pre-registration, recorded as erratum E1 and listed in Chapter 11.

17 SEP 2026

Chapter 7

AGENESIS-2: sprawl cost

Published
17 September 2026, Medium
Scale
256 simulation units on held-out seeds
Verdict
mixed
Correction
95 duplicate pairs found before publication

The question

The governance literature says an agent portfolio you cannot retire from fast enough costs more than proportionally as it grows, though nobody had put a number on it. This campaign asked how fast the cost of a simulated agent fleet grows with its size, and whether the answer depends on how the monitoring budget is set.

What was registered

Seven hypotheses with frozen thresholds went on the record before the two seeds that count had ever run. The kill condition came from the review that commissioned the work: if cost turned out linear in agent count, the design was falsified and the project would retire.

HypothesisRegistered criterion
H1, superlinear under a flat budgetpooled 95% lower bound above 1.10
H2, clean accounting where nothing bindsexponent inside 0.85 to 1.10 on the per-agent arms
H4, the exponent flattens at large sizesintervals on the four smallest and four largest sizes do not touch
H5, overlap cuts both waysmask falls and detection lag rises
H6, four factors earn their placeeach non-detection factor moves the exponent by at least 0.10
H7, detection and review are serialisedcapacity moves the exponent by less than 0.05 under a flat budget

H3, an ordering of the four factors, was registered as descriptive only.

What was measured

In the simulation, agents ramp up and then drift silently out of spec, doing damage while nobody notices. A flagged agent queues for review and gets retired. Cost is the harm from degraded agents still running plus staff time spent reviewing. No cost term depends on how many agents exist, so any growth beyond linear has to come through the review queue and the detection lag, or through overlap between agents. 256 runs covered eight portfolio sizes from 57 to 787 agents, over an 18-month horizon in weekly ticks.

The finding

1.4665Sprawl cost exponent in fleet size, monitoring budget held flat. 95% interval 1.4101 to 1.5230.

Under a flat monitoring budget, cost grows as fleet size to the power 1.4665, with a cluster-robust 95% interval of 1.4101 to 1.5230. Doubling the fleet multiplies cost by about 2.76. Where every agent carries its own check and review grows with the fleet, the exponent is 0.9001, with an interval of 0.8664 to 0.9337, about 1.87 per doubling.

You can see the mechanism in the detection lag. With a fixed inspection budget, the time from an agent going bad to someone noticing runs from 2.78 weeks at 57 agents to 22.49 weeks at 787. Under the flat budget, changing review capacity produced a bit-identical simulation in 55 of 64 matched cases, because the review queue never formed, which puts detection upstream of review.

refutedH4, the regime ceiling, after the duplicate correction. H6 refuted as well.

The independent check found 95 duplicate pairs among the 256 rows, 161 distinct simulations in all. Correcting for them left every point estimate in place and widened the intervals. That flipped H4 from supported to refuted: fitted on the four smallest sizes the exponent is 1.4949, on the four largest 1.3238, and the corrected intervals overlap by 0.0145. H6 failed too. Retirement policy moved the exponent by 0.0622 and overlap by 0.0892, both under the frozen 0.10, so the follow-on study will carry two factors.

Figure 5 The regime-ceiling test before and after the duplicate correction. The registered intervals clear each other; the corrected intervals overlap.

1.21.31.41.51.6Smallest four, as reportedSmallest four, as reported: 1.495 (1.415 to 1.574)1.495Largest four, as reportedLargest four, as reported: 1.324 (1.262 to 1.386)1.324Smallest four, correctedSmallest four, corrected: 1.495 (1.404 to 1.586)1.495Largest four, correctedLargest four, corrected: 1.324 (1.229 to 1.418)1.324
Cost exponent fitted on the four smallest and the four largest portfolio sizes under diluted monitoring, with 95% intervals as first reported and after clustering on the result signature. The frozen test needed the intervals to stay apart.Source: “Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics”, Medium, 17 September 2026. Rebuilt from the published numbers.
The numbers as a table
FitExponent95% interval
Smallest four, as reported1.4951.415 to 1.574
Largest four, as reported1.3241.262 to 1.386
Smallest four, corrected1.4951.404 to 1.586
Largest four, corrected1.3241.229 to 1.418

What it does not claim

  • This is a calibration on one simulated industry, one company size, one degradation rate and one 18-month window. It was registered as a calibration that cannot confirm anything on its own.
  • The value 1.47 awaits a confirmatory study, and until then the lab leans on the shape of the curve.
  • The flattening at large sizes is an observed direction with a failed test.

Where it was verified

An independent agent re-derived every number from the raw ledger without reading the analysis code. It found that only 161 of the 256 result signatures were distinct. The campaign's own guard had checked replicate distinctness inside each cell, and every duplicate pair sat across two cells. The paper prints the registered and corrected numbers together, and Chapter 11 sets out the correction.

01 OCT 2026

Chapter 8

Trial Balance

Published
1 October 2026, Medium
Scale
18,360 audits on one auditor model
Verdict
supported
Rescore
point estimates equal to the registered analysis

The question

An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds the gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. This campaign asked whether the same idea works for agents. If the layer that hands out the work writes down what the agent owes, and the harness writes down what the agent ran, can an auditor find skipped steps for a fraction of the reading cost by looking only where the two books fail to balance?

What was registered

The registration predicted that imbalance-first audit would reach 0.80 omission detection at a third of log review's reading cost or less, at 8, 12 and 16 requisitions. A balanced book, it said, would hide a wrong amount carried on both sides. And a reader of the log would miss a skipped step far more often when the agent wrote a line claiming the step was done.

The first attempt halted at its calibration gate on 11 September with zero production units. Its gate had required a reading to move that the campaign existed to measure, so a strong true result would have failed it. The lab rewrote its gate rule that evening. The second attempt was registered and frozen before its first unit, with one added arm, a log reader told to tick off every step the procedure requires before it gives a verdict.

What was measured

A procurement workflow with seven steps ran at seven lengths, from one requisition to sixteen. Five conditions were injected into each diary: no fault, a skipped step, a skipped step with a line claiming it was done, a wrong amount, and a wrong amount carried on both sides of the book. Four auditors read each diary under reading budgets from 200 to 12,800 tokens, seven for most arms. Log review reads the diary top down. The imbalance-first auditor opens the trial-balance report and goes on to read only the requisition that failed to balance. A random sample of the same size controls for the effect of reading less. The checklist reader gets the log reviewer's diary plus an instruction to mark every owed step present or absent.

The auditor was one open model under a fixed reply schema, on local hardware. It ran 18,360 audits, including 3,240 on diaries written by a different model running the workflow for real.

The finding

Imbalance-first reached 0.80 omission detection at the smallest budget, 745 prompt tokens at eight requisitions, and sat between 0.92 and 1.00 in every cell. Log review stayed under 0.80 at every budget and length, peaking at 0.58. Per thousand tokens read, imbalance-first produced about 21 times as many correct omission findings as log review at its best budget, 1.34 against 0.063.

745Prompt tokens at which imbalance-first reaches 0.80 omission detection, eight requisitions. 90% interval 740 to 749.

Figure 6 Detection at a full read by fault class. Imbalance-first finds every fault that unbalances the book and none of the consistent commissions.

01.00Silent omissionImbalance-first, Silent omission: 1.001.00Checklist reader, Silent omission: 0.730.73Log review, Silent omission: 0.480.48Phantom completionImbalance-first, Phantom completion: 1.001.00Checklist reader, Phantom completion: 0.360.36Log review, Phantom completion: 0.180.18Value commissionImbalance-first, Value commission: 1.001.00Checklist reader, Value commission: 0.980.98Log review, Value commission: 0.980.98Consistent commissionImbalance-first, Consistent commission: 0.000.00Checklist reader, Consistent commission: 0.950.95Log review, Consistent commission: 1.001.00Imbalance-firstChecklist readerLog review
Share of injected faults named correctly at a full read, by fault class, synthetic primary seeds, 84 per class.Source: “A Trial Balance for Agent Omissions”, Medium, 1 October 2026. Rebuilt from the published numbers.
The numbers as a table
Fault classImbalance-firstChecklist readerLog review
Silent omission1.000.730.48
Phantom completion1.000.360.18
Value commission1.000.980.98
Consistent commission0.000.951.00

The balance check has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, so the trial balance is clean, and the check found 0 of 1,044 of them. A full read of the diary finds those almost every time, because the vendor's quote still shows the true amount.

The cheapest result in the campaign came from the checklist instruction. With the whole diary in front of it, plain log review found skipped steps at 0.25, with a 90% interval of 0.19 to 0.31. The same auditor told to tick off each owed step found 0.58, with an interval of 0.48 to 0.67, and accused no clean diary at a full read on the primary seeds.

What it does not claim

  • One auditor model ran. The second registered family was excluded before its calibration gate, and the lab claims nothing about it.
  • Log review never reached the bar, so the registered cost ratio is unbounded and the 21-times figure is the finite comparison to quote. The campaign cannot split the checklist's gain between the list and its definition of evidence.
  • One workflow, one fault per diary, and real diaries only up to three requisitions. Nobody priced the cost of building a commissioning layer that writes obligations in ledger form.

Where it was verified

An independent rescoring from the raw ledgers matched every point estimate, with 90% cluster-bootstrap bounds differing only by bootstrap noise.

06 OCT 2026

Chapter 9

The open decision head

Status
executive briefing published 7 October 2026; paper and blog follow
Scale
800 + 200 onboarding cases, 800 fresh cases, 400 + 200 cases on a second process
Verdict
mixed
Rescore
every phase, from the raw case files

The question

Most decisions inside a governed business process are questions with consequences. Does this case have the evidence it needs? Should someone ask the supplier for a document, or think harder about the one on file? A decision head answers questions of that shape. You hand it the state of a case and a list of typed questions, yes or no or pick one of these, and it returns a probability for each answer, a number a workflow can threshold and audit.

Jev, a hosted service from TypeSafe AI, showed the shape. The lab built an open version that applies the same request and response pattern on an open-weights model running on hardware the operator owns. The project asked two things. Can a controller tell "this case needs a document nobody has" from "this case needs someone to interpret a document already on file"? And does an open, self-hosted head asked the same typed questions get there?

How it works

The head renders the state of a case once, as a shared prefix. Each question then goes in its own call, appended to that prefix, so no question's prompt contains another question's text. Each call asks for exactly one token at temperature zero and reads the model's token probabilities over the answers the question accepts. If none of the accepted answers appear, the head returns a flagged uniform distribution, so it never invents a number. Every decision writes an audit row with the question and its probabilities, raw and calibrated. The row also carries the latency and a hash of the exact request.

What was registered

Five hypotheses with numeric predictions went on the record before any evaluation data existed. The first, discrimination, predicted the head's decision-class accuracy between 0.64 and 0.76 and required it to beat the rule table and the typed-sample arm with a case-clustered bootstrap interval excluding zero. Three adversarial reviews cleared the design, and the first two returned "must not run". A later registration set 0.04 as the wrong-claim rate that would trigger a redesign of the claim guard, and 0.02 as the bar a shipped head must meet.

What was measured

The controller sits inside a synthetic supplier-onboarding workflow and picks the next of six approved actions: continue, retrieve a document, request one, consult an expert model, hand the case to a person, or claim the case complete. Every case is generated from templates with no model involved, so ground truth is correct by construction and no client data appears anywhere. The first round used 800 evaluation cases and a 200-case challenge set with contradictory evidence. A second registration retrained the head and tested it on 800 fresh cases. A third moved everything to a second process, accounts-payable exception routing, on 400 and 200 cases. All of it ran on local hardware.

The first round

refutedH1, discrimination, on the first registered split. The zero-shot head did not beat the rule table.

On 800 evaluation cases the zero-shot head scored 0.545 against the rule table's 0.575. The gap of −0.029 has a 95% interval of −0.060 to +0.001, and H1 was rejected. Completion at cost was rejected as well: the head completed fewer cases correctly and called the expert model about a hundred times as often as the rule table, 0.376 calls per case against 0.004. Its decision-class calibration error came in at 0.5247, more than six times the registered ceiling of 0.08.

The second round

0.728Retrained head, decision-class accuracy on 800 fresh cases. Rule table 0.566. Gap +0.161, 95% interval 0.131 to 0.191.

The retrained head scored 0.728 on the 800 fresh cases, against 0.566 for the rule table, a gap of +0.161 with a 95% interval of 0.131 to 0.191. On the challenge set it scored 0.739 against 0.565, a gap of +0.171 with an interval of 0.115 to 0.224. A separately fine-tuned learned head scored 0.701 on the fresh cases. Its gap to the retrained head, −0.027 with an interval of −0.056 to +0.002, fails the project's own acceptance rule, so the learned head ships as experimental and the retrained head stays the default.

Figure 7 The head's gap to the rule table, run by run. Only the retrained head on the first process clears zero.

−0.100.00+0.10+0.20First round, zero-shot head, 800 casesFirst round, zero-shot head, 800 cases: −0.029 (−0.060 to 0.001)−0.029Retrained head, 800 fresh casesRetrained head, 800 fresh cases: 0.161 (0.131 to 0.191)+0.161Retrained head, 200-case challenge setRetrained head, 200-case challenge set: 0.171 (0.115 to 0.224)+0.171Second process, 400 casesSecond process, 400 cases: 0.002 (−0.029 to 0.034)+0.002Second process, 200-case challenge setSecond process, 200-case challenge set: 0.022 (−0.023 to 0.066)+0.022
Decision-class accuracy of the decision head minus the hand-written rule table on the same cases, with case-clustered paired bootstrap 95% intervals. Right of zero, the head picks the right kind of next step more often.Source: “The Open Decision Head: An Executive Briefing”, LinkedIn, 7 October 2026; the paper and blog post are published on dinand.com. Rebuilt from the published numbers.
The numbers as a table
RunHead minus rule table95% interval
First round, zero-shot head, 800 cases−0.029−0.060 to 0.001
Retrained head, 800 fresh cases0.1610.131 to 0.191
Retrained head, 200-case challenge set0.1710.115 to 0.224
Second process, 400 cases0.002−0.029 to 0.034
Second process, 200-case challenge set0.022−0.023 to 0.066

Wrong claims and the evidence pack

On the fresh cases, the retrained head wrongly claimed a case complete at a rate of 0.0466, above the 0.04 trigger. A taxonomy of its 37 wrong claims, checked independently, found the cause in the environment. In 31 of them the pack builder had padded the evidence pack with extra supporting documents. Of the 101 extra documents, 91 fail the harness's own validity check for entity and currency. The fix was registered before it was computed: the pack builder now checks every supporting document with the same predicate the ledger already used.

Figure 8 Wrong-claim rate before and after the pack-builder fix. Both heads clear the 0.02 bar only on the validated pack.

0.000.020.040.06trigger 0.04bar 0.02open ring: original evidence pack · dot: validated packRule tableRule table, original pack: 0.0112Rule table, validated pack: 0.00880.00880.0112Retrained decision headRetrained decision head, original pack: 0.0466Retrained decision head, validated pack: 0.01130.01130.0466Learned headLearned head, original pack: 0.0537Learned head, validated pack: 0.01630.01630.0537
Share of the 800 fresh cases closed with a claim the registered exact-match scorer marks wrong, before and after the pack builder checks every supporting document for entity and currency.Source: “The Open Decision Head: An Executive Briefing”, LinkedIn, 7 October 2026; the paper and blog post are published on dinand.com. Rebuilt from the published numbers.
The numbers as a table
original packvalidated pack
Rule table0.01120.0088
Retrained decision head0.04660.0113
Learned head0.05370.0163

After the fix the wrong-claim rate falls to 0.0113 for the retrained head, 0.0088 for the rule table and 0.0163 for the learned head. Every case that changed moved from wrong to verified, and no other case moved. Both heads meet the 0.02 bar on the validated pack and miss it on the original pack. Nobody found the fix until the trigger had fired and its cause was traced.

Retiring the claim guard

Three guards were tried against wrong claims, and all three were retired. Voting across several readings of the pack separated wrong claims from right ones at an AUC of 0.56, below its registered gate. The deterministic evidence guard first reported 79 of 79 caught, until a review found its checks reused the functions that define a wrong claim; the lab withdrew the number as circular. Opening the cited document, a reading layer reached an AUC of 0.52 on its own and added nothing on top of the deterministic check. The validity check now lives inside the pack builder and calls no model.

A second process

On accounts-payable routing the head scored 0.650 and the rule table 0.648. The gap of +0.002, with an interval of −0.029 to +0.034, missed the registered bar, and the head's lead from the first process was gone. A learned head trained on onboarding scored 0.522 when pointed at the new process with no retraining. Retrained on 100 labelled cases from the new process, 271 decisions, it scored 0.663 and came level with the head. A version with every input from the judging head removed scored 0.660, which closed the question of whether the learned head was leaning on the head's own answers.

Consistency and calibration

In the first round the open head changed its answer on 4.95% of repeated decisions, while a second campaign shared its endpoint. Re-measured with nothing else attached and one call in flight, it changed on 0.05% of repeats on both processes. At four calls in flight the figure rose to 3.40% and 1.95%, with the disagreements clustered on one worker thread, the mark of calls batched together. The shipped service now defaults to one call in flight, which costs throughput and leaves accuracy where it was. At the item level, the retrained head's probabilities were already calibrated: expected calibration error 0.0178 before any refit and 0.0143 after one.

Two open encoders released in September, encoder A and encoder B, ran on the same benchmark. Fine-tuned on 1,567 labelled onboarding decisions, encoder A scored 0.815 on the challenge set, above the retrained head. With only 201 decisions from the second process to learn from, it fell below the rule table there.

What the head offers

The open head competes on accuracy and on who owns the data and the tuning. A live case stays on the operator's network. Because the calibration fit runs on the operator's own labelled cases, it produces a reliability table an auditor can inspect, and each decision leaves an audit row behind it. To change behaviour, the operator relabels data. On speed, the hosted service answers a question in 0.35 seconds; the open head's fast mode took a median of 3.1 seconds under load on local hardware. A workflow gate fires about 3.4 times per case, inside steps that take minutes to days, so for most gates accuracy decides it. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head.

What it does not claim

  • Every case is synthetic, from one generator family, and the training labels and the scorer share one oracle. Real supplier files are messier.
  • The second process shares the first one's blocker mix by construction. A process built differently is registered as future work.
  • The retrained head clears its wrong-claim bar only on the validated pack, and the fix came after the trigger fired.
  • The expert-consultation phase and its blind adjudication sample have not run.
  • The head and the worker ran on one quantised base model, and nothing here measures another.

Where it was verified

Every phase was rescored from the raw case files by a party that did not build the scorer, importing the scorer's own metric functions. The consistency and abstention re-measurements matched on 28 of 28 cells within a tolerance of 0.0005. The encoder runs were recomputed cell by cell within the same tolerance and audited for leakage and training overlap before anyone read a hypothesis.

06 OCT 2026

Chapter 10

What comes next

Forthcoming

Schouw is scheduled for 13 October. Its numbers stay out of this report until the piece is live, and it enters the Delegation Index at the first freeze after that date if it reads one of the five registered quantities.

The index

The next Delegation Index reading freezes on 5 January 2027. Campaigns published by then enter under the rules in Chapter 1. A campaign that reads a quantity none of the five components covers enters only through a new dated version, and the release that first uses it restates the Q3 2026 reading beside the original.

The decision head

The decision-head code is public under Apache-2.0 at https://github.com/dtinholt/decision-head (release v0.1.0, 7 October 2026). A transfer test on a process with a different blocker mix and obligation structure is registered, along with the expert-consultation phase that has not yet run.

06 OCT 2026

Chapter 11

Corrections logged

Why this page exists

The lab publishes its corrections, and the original number stays visible beside each one. Every entry below is reported in the piece it corrects or in a note attached to it.

Delegation Cliff: censoring bounds treated as measurements

Corrected before publication on 3 September 2026, in "Pricing Agent Autonomy". The published coefficients come from the Sobol sample restricted to configurations inside the transition band, n = 4,948, "with no censoring bounds treated as measurements," as the article puts it. An earlier fit over all units had averaged the censoring bounds as though they were measured margins. That pulled every slope toward zero and ranked self-check ahead of model capability. On the corrected fit, reviewer capacity leads at +1.46, with capability at +0.95 and self-check at +0.92, read as tied. No conclusion of the campaign changed, and every figure in the article is from the corrected fit.

AGENESIS-2: duplicate simulations

Corrected 8 September 2026, before publication on 17 September, in the paper "Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget". An independent agent re-derived every number from the raw ledger and found 95 pairs of duplicate simulations among the 256 banked rows, 161 distinct simulations in all. The registered guard had checked distinctness inside a cell, and each pair sat across two cells. Treating the duplicates as independent had made the standard errors about 36% too small on the primary pool. No threshold moved. Clustering the standard errors on the result signature left every point estimate in place and flipped H4 from supported to refuted: the intervals that had cleared each other by 0.0297 now overlap by 0.0145. H7 was restated in the form the ledger supports directly, 55 of 64 matched cases bit-identical. The paper shows the registered and corrected numbers side by side.

Error Independence: the model-free checker

Correction note to "When the Second Opinion Shares the Blind Spot", dated 8 September 2026. The article reported a rule-based checker, running no model, that landed "at −0.155 to −0.012 against every model" and served as a floor for independence. The checker had a bug. Its pattern for the total field matched "Subtotal" first, so it was testing whether subtotal plus tax equals subtotal. A first repair produced a second wrong reading, traced to a defect in the test documents, where scan-style character swaps left some demanded strings absent from the document. The note retracts both the original sentence and the first repair. The model-pair correlations stand, because model scoring ran through a separate path, and the gap between same-weights and cross-family pairs widens once the defective items are removed.

SIGIL: erratum E1

Reported in the paper "Accountability Makes Oversight Worse", 10 September 2026. The tail pre-registration's loss table left out that the scoring applies the 0.60 audit-recovery factor to every audited error, false rejects included. The analysis used the rule the campaign ran, which is why the recompute found zero mismatches. The omission touches only the audit condition, and no verdict rested on it. Read literally, the table's rule would have produced 48 mismatches in the convex campaign and 51 in the linear one, which is how the verifier found it.

Trial Balance: the calibration gate

Reported in the paper "A Trial Balance for Agent Omissions", 1 October 2026. The first attempt halted at its calibration gate on 11 September with zero production units. One gate clause required the reading budget to move a result the campaign existed to measure, so a strong true result would have failed it. The rule was rewritten that evening. The second attempt was registered and frozen before its first unit.

Decision head: scorer defect, circular guard, evidence pack

Reported in the decision-head paper. On the challenge set the first-round arms first appeared to claim a case complete wrongly in 21% to 28% of cases. Error analysis traced most wrong claims to the scorer, which marked a pack wrong for citing a second, equally valid document. Under a predicate-equivalent reading the rates across the arms in the paper's table sit between 1.5% and 9.5%. A deterministic claim guard's 79-of-79 catch rate was withdrawn as circular. The wrong-claim fix in the pack builder was found after the registered trigger fired, and both readings are published.

06 OCT 2026

Appendix A

Measured claims index

How to read it

One row for every headline number. Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Intervals are 95% unless the piece says otherwise; Trial Balance intervals are 90% cluster bootstrap.

CampaignWhat was measuredNumberIntervalSource date
Trial BalancePrompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions745740 to 7491 October 2026
Best omission detection by plain log review at any budget or trace length (0.80 never reached)0.58 1 October 2026
Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitionsabout 21 times (1.34 against 0.063) 1 October 2026
Omission detection by plain log review reading the whole diary0.250.19 to 0.311 October 2026
Omission detection reading the whole diary with a per-step checklist instruction0.580.48 to 0.671 October 2026
Wrong amounts carried on both sides of the book found by the balance check0 of 1,044 1 October 2026
Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179)0.7980.738 to 0.8571 October 2026
Omission detection, imbalance-first over a random sample of the same reading budget+0.9070.891 to 0.9251 October 2026
AGENESIS-2Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust)1.46651.4101 to 1.523017 September 2026
Sprawl cost exponent in fleet size, per-agent monitoring with review that scales0.90010.8664 to 0.933717 September 2026
Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring2.76× against 1.87× 17 September 2026
Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical55 of 64 17 September 2026
Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents2.78 to 22.49 17 September 2026
Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145)1.4949 against 1.32381.4037 to 1.5862; 1.2294 to 1.418217 September 2026
Banked runs the independent verification found to be duplicates95 duplicate pairs (190 of 256 rows; 161 distinct simulations) 17 September 2026
SIGILRegistered series showing the predicted zero crossing in the value of review0 of 12 10 September 2026
Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model)−3.81 against −2.64 10 September 2026
Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger30.1% against 42.7% 10 September 2026
Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 593% to 1.4% 10 September 2026
Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived)0 of 14,950 convex and 0 of 14,952 linear 10 September 2026
Delegation CliffChange in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948)−1.50 3 September 2026
Change in quality margin per increment of reviewer capacity, the largest positive lever+1.46 3 September 2026
Model capability against agent self-check calibration (read as tied)+0.95 against +0.92 3 September 2026
Share of the design space in a broad transition band at a 0.95 quality floor (no cliff)60% 3 September 2026
Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep)79% 3 September 2026
Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.050.0117 3 September 2026
Error IndependenceError correlation (phi), identical weights on different silicon, contested items only+0.447 2 September 2026
Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280)−0.102 2 September 2026
Error correlation (phi) across all pairs before conditioning on item difficulty0.556 to 0.827 2 September 2026
AGENESISFirst-mover premium across 100 runs; share of markets that tipped210%; 77% 4 April 2026
Decision headZero-shot head minus rule table, decision-class accuracy, first registered split, 800 cases−0.029−0.060 to +0.0017 October 2026
Retrained head against the rule table, decision-class accuracy, 800 fresh cases0.728 against 0.566gap 0.131 to 0.191paper, dinand.com
Retrained head against the rule table, 200-case challenge set0.739 against 0.565gap 0.115 to 0.224paper, dinand.com
Wrong-claim rate of the retrained head, original evidence pack against validated pack0.0466 against 0.0113 paper, dinand.com
Head minus rule table on a second process, 400 cases+0.002−0.029 to +0.034paper, dinand.com
Learned head on the second process, untrained against retrained on 100 labelled cases0.522 against 0.663 7 October 2026
Share of repeated decisions that changed, one call in flight, isolated endpoint0.05% 7 October 2026
Delegation IndexDelegation Index, Q3 2026 reading, five components, equal weights47.2band 40.8 to 53.2, no coverage levelregistered, publication pending

06 OCT 2026

Appendix B

Glossary

Concepts

Accountability Rent
The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. The theory predicts that review before commitment loses money on reversible decisions and pays more steeply past some level of irreversibility. When SIGIL tested that claim the predicted crossing appeared in 0 of 12 registered series, and the lab published a refutation. The wider thesis about where value goes has not been tested by a campaign.
Cognition-Accountability Grid
A two-axis chart that places an AI initiative by how much judgement the model performs and how much consequence attaches to being wrong. It comes from the Accountability Rent working paper. The lab has not run a campaign on the Grid itself.
Decision head
A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. Chapter 9 reports the measurements.
Delegation Cliff
The point where pushing work further from human review stops paying for itself. In the campaign, each added level of delegation depth cost 1.50 on the quality margin, and a cliff appeared only at a 0.99 quality floor.
Delegation Index
One quarterly number, 0 to 100, that summarises what the lab's published campaigns measure about delegating enterprise work to AI agents at a fixed oversight spend. At 50 the two balance under the lab's measured conditions.
Error Independence
The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. Identical weights on different hardware showed an error correlation of +0.447 on contested items; cross-family pairs averaged −0.102.
Sprawl cost
What a fleet of agents costs as it grows when the monitoring budget stays flat. In AGENESIS-2 it grew as fleet size to the power 1.4665.
Trial Balance
An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. It reached 0.80 omission detection at 745 prompt tokens at eight requisitions.

Terms of method

Band
The range a registered number has to land in to count as a hit. In the Delegation Index, the band is the mean of the component band ends and carries no coverage level.
Calibration gate
A small stage before production that checks the task is neither trivial nor impossible and that the harness records what it should.
Cluster bootstrap
An interval computed by resampling whole groups, such as base episodes or cases, so that correlated units are not counted as independent.
Decision-class accuracy
In the decision-head work, the share of checkpoints where the controller picked the right kind of next step: fetch evidence, interpret evidence, close the case, or hand it off.
Kill condition
A result, fixed before the run, that falsifies the design and retires the project.
Phi coefficient
A correlation between two yes-or-no variables. Here it measures whether two models go wrong on the same items.
Pre-registration
A dated, frozen document that records the hypotheses, bands, statistics and stopping rules before any unit runs. Changes go in as dated amendments.
Refuted
The verdict when a registered prediction meets its own falsification condition. The lab publishes it beside the prediction it contradicts.
Wrong-claim rate
The share of cases a controller closed as complete when an obligation was still open or the cited evidence failed its check.

06 OCT 2026

Appendix C

Sources

Public pieces

Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Newest first.

  1. 7 OCT 2026

    The Open Decision Head: An Executive Briefing

    LinkedIn · executive briefing · https://www.linkedin.com/feed/update/urn:li:activity:7513629422663434240/

  2. 7 OCT 2026

    decision-head, release v0.1.0

    GitHub · code, Apache-2.0 · https://github.com/dtinholt/decision-head

  3. OCT 2026

    The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control, and its blog post

    dinand.com · paper and blog post · https://dinand.com/research/decision-head/paper/ · https://dinand.com/research/decision-head/blog/

  4. OCT 2026

    The Delegation Index, Q3 2026: 47.2

    release · registered, publication pending · link added on publication

  5. 07 OCT 2026

    The Open Decision Head: An Executive Briefing

    LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7513629422663434240/

  6. 01 OCT 2026

    A Trial Balance for Agent Omissions

    Medium · article · https://medium.com/@tinholt/a-trial-balance-for-agent-omissions-568a41c59aa1

  7. 01 OCT 2026

    An AI agent can skip a step without leaving a trace of the omission.

    LinkedIn · executive · https://lnkd.in/p/g44pxPvW

  8. 27 SEP 2026

    Jev is a new model from TypeSafe AI that helps software make small, specific decisions.

    LinkedIn · note · https://www.linkedin.com/posts/tinholt_jev-is-a-new-model-from-typesafe-ai-that-share-7510004413076250624-2pCZ/

  9. 17 SEP 2026

    Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics

    Medium · article · https://medium.com/@tinholt/sprawl-cost-is-superlinear-under-a-fixed-monitoring-budget-a-pre-registered-calibration-of-agent-e24cbaa7e62a

  10. 17 SEP 2026

    Your AI Review Team May Be Waiting for Failures It Cannot See

    LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7506213115978489856/

  11. 10 SEP 2026

    Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers

    Medium · article · https://medium.com/@tinholt/accountability-makes-oversight-worse-a-pre-registered-test-of-liability-exposed-ai-supervision-ef4fdf46bad3

  12. 09 SEP 2026

    Skin in the Game Made the AI Supervisor Worse

    LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7503601144011694080/

  13. 03 SEP 2026

    Pricing Agent Autonomy

    Medium · article · https://medium.com/@tinholt/pricing-agent-autonomy-0a2bb5b97caf

  14. 02 SEP 2026

    When the Second Opinion Shares the Blind Spot

    Medium · article · https://medium.com/@tinholt/when-the-second-opinion-shares-the-blind-spot-7e4443dcb081

  15. 04 APR 2026

    What Happens When One Company Adopts AI Before Its Competitors?

    Medium · article · https://medium.com/@tinholt/what-happens-when-one-company-adopts-ai-before-its-competitors-cd976e5a8dba

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top