Who Checks the Agents walks through the lab’s research on AI delegation and oversight, chapter by chapter. It opens on a supplier onboarding case where an agent skipped a step its log never recorded, and then goes through six published studies before it ends on the first Delegation Index reading. Every figure on screen is listed below with the piece it comes from.
Chapters
Select a chapter to jump to it in the film.
The numbers on screen
Each figure the film shows, chapter by chapter, with its source. The campaign pages carry the intervals and the limits in full.
1. The question
No figures in this chapter. The onboarding case on screen is illustrative.
2. The control plane
No figures in this chapter. The control plane is defined in plain words.
3. How the lab works
| 95 | duplicate pairs found by the independent pass in AGENESIS-2 (190 of 256 rows; 161 distinct simulations) Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget, 17 Sep 2026 |
| 30 Aug 2026 | SIGIL output contract frozen, before any harness code existed The Measured Enterprise 2026, chapter 6 |
| 2% → 1.4% | SIGIL tier-5 share of errors that were wrong approvals; the executive piece printed 2%, the paper gives 1 of 69. Correction logged 6 Oct 2026 Measured claims index, row sigil-fa-share |
Source: The Measured Enterprise 2026, chapters 2 and 11
4. Delegation Cliff
| 90,880 | simulated agent organisations, zero error units Pricing Agent Autonomy |
| −1.499 | change in quality margin per increment of delegation depth band-only Sobol fit, n = 4,948; descriptive slopes |
| +1.456 | change in quality margin per increment of reviewer capacity same fit |
| 60% | of the design space in a broad transition band at a 0.95 quality floor; the registered cliff prediction failed there replicated on the 8,192-unit Sobol set |
| 79% | of the space fragile at a 0.99 floor, where a cliff appears one 256-unit calibration sweep |
Source: Pricing Agent Autonomy · Medium · 3 Sep 2026
5. SIGIL
| 30,000 | supervised episodes across five irreversibility tiers two full campaigns, one per cost model, three seeds per cell |
| 0 of 12 | registered series showing the predicted gain from review on hard-to-undo decisions under both cost models |
| 30.1% → 42.7% | tier-5 decisions wrongfully blocked, plain review against review with a liability ledger convex cost model |
| 93% → 1.4% | share of worker errors that were wrong approvals, tier 1 to tier 5 (1 of 69) the only errors a blocking reviewer can fix |
Source: Accountability Makes Oversight Worse · Medium · 10 Sep 2026
6. Error Independence
| 150 | invoice-extraction items with generated ground truth, four model setups no model grades another |
| 0.827 | error correlation, identical weights on different hardware, all items (on screen only) before conditioning on difficulty |
| 0.447 | same pair, on the 47 items where the models disagreed the highest in the matrix; one pair |
| −0.102 | mean across the five cross-family pairs, contested items spread −0.469 to +0.280, so the mean describes them poorly |
Source: When the Second Opinion Shares the Blind Spot · Medium · 2 Sep 2026
7. AGENESIS-2
| 1.4665 | exponent of cost in fleet size, monitoring budget held flat 95% interval 1.4101 to 1.5230; 256 runs, 57 to 787 agents |
| ×2.76 | cost multiple per doubling of the fleet under a flat budget executive piece |
| 0.9001 | exponent with per-agent checks, where spend rises with every agent 95% interval 0.8664 to 0.9337 |
| 2.78 → 22.49 | weeks from an agent going bad to detection, 57 to 787 agents, fixed inspection budget two published endpoints |
| H4 refuted | after correcting for 95 duplicate pairs; a calibration on one simulated industry intervals now overlap by 0.0145 |
Source: Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget · Medium · 17 Sep 2026
8. Trial Balance
| 18,360 | audits on one auditor model, one seven-step procurement workflow 3,240 on diaries written by a different model running the workflow |
| 745 | prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions 90% interval 740 to 749 |
| 0.92–1.00 | imbalance-first detection across every cell |
| 0.58 | best omission detection by plain log review at any budget (0.80 never reached) |
| 0 of 1,044 | wrong amounts carried on both sides of the book found by the balance check the registered blind spot; a full log read finds most of them |
Source: A Trial Balance for Agent Omissions · Medium · 1 Oct 2026
9. The decision head
| 0.545 / 0.575 | first round, open head against the rule table, 800 cases: did not beat it, within noise (on screen only) gap −0.029, interval −0.060 to +0.001 |
| +0.161 | retrained head 0.728 against the rule table's 0.566, 800 fresh cases 95% interval 0.131 to 0.191 |
| 0.047 → 0.0113 | wrong-claim rate, retrained head, before and after the evidence-pack fix the fix came after the 0.04 trigger fired; both readings published |
| 0.35 s / 3.1 s | hosted service, single question, against the open head's median under four parallel streams |
| 0.650 / 0.648 | second process, open head against the rule table: level two synthetic processes built the same way; no general claim that owned hardware beats a hosted model |
Source: The Decision Head · dinand.com · 7 Oct 2026
10. The Delegation Index
| 47.2 | Delegation Index, Q3 2026 band 40.8 to 53.2; the band contains 50 |
| 26.7 | sprawl component, the lowest of five (delegation price 49.3, oversight yield 46.9, audit detection 58.3, reviewer independence 55.1) |
| 20,154 | units behind the five quoted numbers, of different kinds the lab's own synthetic measurements; market and firm data lie outside its scope |
Source: The Measured Enterprise 2026 · dinand.com · 7 Oct 2026
How the lab works
The film rests on four rules behind every number. The full eight-step protocol is on the methods page.
- Registered first
- Each prediction is written down, with its band and its kill condition, then dated and frozen before any run.
- Checked from raw files
- An independent pass with fresh code recomputes the published numbers from the raw records.
- Published either way
- Refuted predictions appear beside the prediction they contradicted, and corrections keep the original number visible.
- Owned hardware
- Runs use open models on the lab’s own machines. The decision-head study added one hosted service, reached through a broker.
Transcript
The spoken text in full. The paragraph being spoken is marked while the film plays, and each timestamp jumps to its paragraph.
1. The question
A supplier onboarding case is about to close. One document is still missing. The agent skipped the step that would have caught it, and its log has no line for that step.
Organisations now hand decisions like this one to AI agents. Each deployment carries a bet about cost, as work moves further from human review.
Dinand Tinholt runs a small research lab that measures that cost.
2. The control plane
Engineers call one layer of an agent system the control plane. It sits between people and their agents and sets what an agent may do. After the work, it checks the result and records both steps.
A reviewer with the power to block is one form of the layer. A second model checking the first is another, and the lab's campaigns measure the cost and catch rate of each.
3. How the lab works
Before a campaign runs, the lab writes down its prediction with the band each number has to land in. A second line names the result that would end the work, and the document is dated and frozen.
The runs use open models on hardware the lab owns. The decision-head study added one hosted service, reached through a broker. After a run, an independent pass with fresh code recomputes the numbers from the raw files.
A refuted prediction is published next to the prediction it contradicted. When a number is corrected, the original stays visible beside it.
4. Delegation Cliff
Delegation Cliff looks for the point where delegation stops paying.
A simulation scored 90,880 agent organisations on their margin, the gap between quality delivered and the floor required. Delegation depth means how far work sits from human review. The fit used the 4,948 cases inside the transition band, and its slopes are descriptive. Each step of depth cost 1.499. Reviewer capacity added 1.456 per unit.
At a quality floor of 0.95, 60 percent of the space sits in a broad transition band. The registered cliff prediction failed there. One sweep of 256 units found a cliff at a floor of 0.99.
5. SIGIL
Accountability for an AI reviewer was the subject of SIGIL.
The campaign ran 30,000 supervised episodes. An AI worker made synthetic business decisions across five tiers. The tiers ran from reversible invoices to permanent public statements. A second copy of the same model reviewed it, with power to block. In one design, the reviewer carried a ledger of the harm its past approvals caused.
The registered prediction, that review pays most on decisions hardest to undo, failed in all 12 series. At the top tier, plain review wrongly blocked 30.1 percent of decisions. The ledger charged nothing for a refusal and raised wrongful blocking to 42.7 percent.
A blocking reviewer fixes one kind of error, a wrong approval. Wrong approvals made up 93 percent of the worker's errors at the lowest tier and 1.4 percent at the highest. One model family sat in both seats, and the ledger was a thin, simulated form of accountability.
6. Error Independence
A common resilience design runs the same model again on another machine and uses its answer as a check. Error Independence put that design to a test.
The lab gave 150 invoices with known answers to four model setups. One pair ran identical weights on different hardware. Correlation near one means two models fail on the same items. On the 47 items where the models disagreed, that pair's errors correlated at 0.447.
Pairs from different model families averaged minus 0.102. That mean hides a spread from minus 0.469 to plus 0.280. The same-weights reading comes from one pair on one task.
7. AGENESIS-2
AGENESIS-2 measured fleet cost with the monitoring budget held flat.
The simulation covered 256 runs. Fleets ranged from 57 to 787 agents. With monitoring flat, cost grew as fleet size to the power 1.4665. Each doubling multiplied cost by about 2.76. Giving every agent its own check brought the exponent to 0.9001, with spend rising for every agent added.
Under a fixed inspection budget, the wait before anyone noticed a bad agent went from under three weeks to over twenty-two. The work is a calibration on one simulated industry. Verification found duplicate runs and moved one verdict to refuted.
8. Trial Balance
Trial Balance borrowed the second entry from bookkeeping.
The layer that assigns the work keeps one book of what the agent owes. The harness keeps a second book of what the agent ran, and the auditor reads only where the two fail to balance. The test ran 18,360 audits of a seven-step procurement workflow. At eight requisitions, the method found 80 percent of skipped steps. That took 745 tokens of reading.
Detection stayed between 0.92 and 1.00 in every cell of the test. The best reading for plain log review was 58 percent.
A wrong amount entered before the work is assigned sits in both books. The balance check found none of the 1,044 faults of that kind, as the registration predicted. A full read of the log catches most of them.
These results come from one auditor model on one workflow. Each diary carried one fault, and real diaries ran up to three requisitions, with the ledger layer's cost left unpriced.
9. The decision head
The lab also built a decision head, a small service a company runs on its own hardware. A workflow asks it typed questions about a case. It returns a probability for each answer and writes an audit row.
In the first round, the head did not beat a hand-written rule table, and the difference sat within noise. After retraining, the head scored 0.728 on fresh cases. The rule table reached 0.566 on the same cases. The lead was 0.161. Its interval runs from 0.131 to 0.191.
The retrained head closed cases wrongly at a rate of 0.047, above the bar set in advance. The cause was in the evidence pack the test environment built. A fix came only after the bar was missed. With it the rate was 0.0113, and both readings are published.
Jev, a hosted service, answered a single question in 0.35 seconds. Under four parallel streams, the open head's median response was 3.1 seconds.
On a second business process, the head's lead was gone. The code is public under the Apache 2.0 licence. The finding covers one decision task on two synthetic processes built the same way.
10. The Delegation Index
Five of the campaigns feed one quarterly number, the Delegation Index. Its scale runs from zero to a hundred. A reading of fifty means the cost of handing work to agents equals what oversight returns at a fixed budget. The first reading is 47.2. The band runs from 40.8 to 53.2.
The reading rests on 20,154 units behind the five numbers, of different kinds. The sprawl component scores lowest, at 26.7. The band contains fifty, so this quarter's evidence fits balance and imbalance alike. The index covers the lab's own synthetic measurements, and market and firm data lie outside its scope.
Ask for the registered prediction, with its date. Which band did each number have to land in? Someone must name the result that would have stopped the work. The figures need an independent recompute from the raw files, using fresh code. Request the negative results as well.
The campaigns and their corrections are at dinand.com slash research.