8 min 54 s, 1080p, English captions. Narration by a locally run neural voice; the music was synthesised for this film, with no samples.

Who Checks the Agents walks through the lab’s research on AI delegation and oversight, chapter by chapter. It opens on a supplier onboarding case where an agent skipped a step its log never recorded, and then goes through six published studies before it ends on the first Delegation Index reading. Every figure on screen is listed below with the piece it comes from.

Chapters

Select a chapter to jump to it in the film.

The numbers on screen

Each figure the film shows, chapter by chapter, with its source. The campaign pages carry the intervals and the limits in full.

1. The question

No figures in this chapter. The onboarding case on screen is illustrative.

2. The control plane

No figures in this chapter. The control plane is defined in plain words.

3. How the lab works

95duplicate pairs found by the independent pass in AGENESIS-2 (190 of 256 rows; 161 distinct simulations)
Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget, 17 Sep 2026
30 Aug 2026SIGIL output contract frozen, before any harness code existed
The Measured Enterprise 2026, chapter 6
2% → 1.4%SIGIL tier-5 share of errors that were wrong approvals; the executive piece printed 2%, the paper gives 1 of 69. Correction logged 6 Oct 2026
Measured claims index, row sigil-fa-share

Source: The Measured Enterprise 2026, chapters 2 and 11

4. Delegation Cliff

90,880simulated agent organisations, zero error units
Pricing Agent Autonomy
−1.499change in quality margin per increment of delegation depth
band-only Sobol fit, n = 4,948; descriptive slopes
+1.456change in quality margin per increment of reviewer capacity
same fit
60%of the design space in a broad transition band at a 0.95 quality floor; the registered cliff prediction failed there
replicated on the 8,192-unit Sobol set
79%of the space fragile at a 0.99 floor, where a cliff appears
one 256-unit calibration sweep

Source: Pricing Agent Autonomy · Medium · 3 Sep 2026

5. SIGIL

30,000supervised episodes across five irreversibility tiers
two full campaigns, one per cost model, three seeds per cell
0 of 12registered series showing the predicted gain from review on hard-to-undo decisions
under both cost models
30.1% → 42.7%tier-5 decisions wrongfully blocked, plain review against review with a liability ledger
convex cost model
93% → 1.4%share of worker errors that were wrong approvals, tier 1 to tier 5 (1 of 69)
the only errors a blocking reviewer can fix

Source: Accountability Makes Oversight Worse · Medium · 10 Sep 2026

6. Error Independence

150invoice-extraction items with generated ground truth, four model setups
no model grades another
0.827error correlation, identical weights on different hardware, all items (on screen only)
before conditioning on difficulty
0.447same pair, on the 47 items where the models disagreed
the highest in the matrix; one pair
−0.102mean across the five cross-family pairs, contested items
spread −0.469 to +0.280, so the mean describes them poorly

Source: When the Second Opinion Shares the Blind Spot · Medium · 2 Sep 2026

7. AGENESIS-2

1.4665exponent of cost in fleet size, monitoring budget held flat
95% interval 1.4101 to 1.5230; 256 runs, 57 to 787 agents
×2.76cost multiple per doubling of the fleet under a flat budget
executive piece
0.9001exponent with per-agent checks, where spend rises with every agent
95% interval 0.8664 to 0.9337
2.78 → 22.49weeks from an agent going bad to detection, 57 to 787 agents, fixed inspection budget
two published endpoints
H4 refutedafter correcting for 95 duplicate pairs; a calibration on one simulated industry
intervals now overlap by 0.0145

Source: Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget · Medium · 17 Sep 2026

8. Trial Balance

18,360audits on one auditor model, one seven-step procurement workflow
3,240 on diaries written by a different model running the workflow
745prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions
90% interval 740 to 749
0.92–1.00imbalance-first detection across every cell
0.58best omission detection by plain log review at any budget (0.80 never reached)
0 of 1,044wrong amounts carried on both sides of the book found by the balance check
the registered blind spot; a full log read finds most of them

Source: A Trial Balance for Agent Omissions · Medium · 1 Oct 2026

9. The decision head

0.545 / 0.575first round, open head against the rule table, 800 cases: did not beat it, within noise (on screen only)
gap −0.029, interval −0.060 to +0.001
+0.161retrained head 0.728 against the rule table's 0.566, 800 fresh cases
95% interval 0.131 to 0.191
0.047 → 0.0113wrong-claim rate, retrained head, before and after the evidence-pack fix
the fix came after the 0.04 trigger fired; both readings published
0.35 s / 3.1 shosted service, single question, against the open head's median under four parallel streams
0.650 / 0.648second process, open head against the rule table: level
two synthetic processes built the same way; no general claim that owned hardware beats a hosted model

Source: The Decision Head · dinand.com · 7 Oct 2026

10. The Delegation Index

47.2Delegation Index, Q3 2026
band 40.8 to 53.2; the band contains 50
26.7sprawl component, the lowest of five (delegation price 49.3, oversight yield 46.9, audit detection 58.3, reviewer independence 55.1)
20,154units behind the five quoted numbers, of different kinds
the lab's own synthetic measurements; market and firm data lie outside its scope

Source: The Measured Enterprise 2026 · dinand.com · 7 Oct 2026

How the lab works

The film rests on four rules behind every number. The full eight-step protocol is on the methods page.

Registered first
Each prediction is written down, with its band and its kill condition, then dated and frozen before any run.
Checked from raw files
An independent pass with fresh code recomputes the published numbers from the raw records.
Published either way
Refuted predictions appear beside the prediction they contradicted, and corrections keep the original number visible.
Owned hardware
Runs use open models on the lab’s own machines. The decision-head study added one hosted service, reached through a broker.

Transcript

The spoken text in full. The paragraph being spoken is marked while the film plays, and each timestamp jumps to its paragraph.

1. The question

A supplier onboarding case is about to close. One document is still missing. The agent skipped the step that would have caught it, and its log has no line for that step.

Organisations now hand decisions like this one to AI agents. Each deployment carries a bet about cost, as work moves further from human review.

Dinand Tinholt runs a small research lab that measures that cost.

2. The control plane

Engineers call one layer of an agent system the control plane. It sits between people and their agents and sets what an agent may do. After the work, it checks the result and records both steps.

A reviewer with the power to block is one form of the layer. A second model checking the first is another, and the lab's campaigns measure the cost and catch rate of each.

3. How the lab works

Before a campaign runs, the lab writes down its prediction with the band each number has to land in. A second line names the result that would end the work, and the document is dated and frozen.

The runs use open models on hardware the lab owns. The decision-head study added one hosted service, reached through a broker. After a run, an independent pass with fresh code recomputes the numbers from the raw files.

A refuted prediction is published next to the prediction it contradicted. When a number is corrected, the original stays visible beside it.

4. Delegation Cliff

Delegation Cliff looks for the point where delegation stops paying.

A simulation scored 90,880 agent organisations on their margin, the gap between quality delivered and the floor required. Delegation depth means how far work sits from human review. The fit used the 4,948 cases inside the transition band, and its slopes are descriptive. Each step of depth cost 1.499. Reviewer capacity added 1.456 per unit.

At a quality floor of 0.95, 60 percent of the space sits in a broad transition band. The registered cliff prediction failed there. One sweep of 256 units found a cliff at a floor of 0.99.

5. SIGIL

Accountability for an AI reviewer was the subject of SIGIL.

The campaign ran 30,000 supervised episodes. An AI worker made synthetic business decisions across five tiers. The tiers ran from reversible invoices to permanent public statements. A second copy of the same model reviewed it, with power to block. In one design, the reviewer carried a ledger of the harm its past approvals caused.

The registered prediction, that review pays most on decisions hardest to undo, failed in all 12 series. At the top tier, plain review wrongly blocked 30.1 percent of decisions. The ledger charged nothing for a refusal and raised wrongful blocking to 42.7 percent.

A blocking reviewer fixes one kind of error, a wrong approval. Wrong approvals made up 93 percent of the worker's errors at the lowest tier and 1.4 percent at the highest. One model family sat in both seats, and the ledger was a thin, simulated form of accountability.

6. Error Independence

A common resilience design runs the same model again on another machine and uses its answer as a check. Error Independence put that design to a test.

The lab gave 150 invoices with known answers to four model setups. One pair ran identical weights on different hardware. Correlation near one means two models fail on the same items. On the 47 items where the models disagreed, that pair's errors correlated at 0.447.

Pairs from different model families averaged minus 0.102. That mean hides a spread from minus 0.469 to plus 0.280. The same-weights reading comes from one pair on one task.

7. AGENESIS-2

AGENESIS-2 measured fleet cost with the monitoring budget held flat.

The simulation covered 256 runs. Fleets ranged from 57 to 787 agents. With monitoring flat, cost grew as fleet size to the power 1.4665. Each doubling multiplied cost by about 2.76. Giving every agent its own check brought the exponent to 0.9001, with spend rising for every agent added.

Under a fixed inspection budget, the wait before anyone noticed a bad agent went from under three weeks to over twenty-two. The work is a calibration on one simulated industry. Verification found duplicate runs and moved one verdict to refuted.

8. Trial Balance

Trial Balance borrowed the second entry from bookkeeping.

The layer that assigns the work keeps one book of what the agent owes. The harness keeps a second book of what the agent ran, and the auditor reads only where the two fail to balance. The test ran 18,360 audits of a seven-step procurement workflow. At eight requisitions, the method found 80 percent of skipped steps. That took 745 tokens of reading.

Detection stayed between 0.92 and 1.00 in every cell of the test. The best reading for plain log review was 58 percent.

A wrong amount entered before the work is assigned sits in both books. The balance check found none of the 1,044 faults of that kind, as the registration predicted. A full read of the log catches most of them.

These results come from one auditor model on one workflow. Each diary carried one fault, and real diaries ran up to three requisitions, with the ledger layer's cost left unpriced.

9. The decision head

The lab also built a decision head, a small service a company runs on its own hardware. A workflow asks it typed questions about a case. It returns a probability for each answer and writes an audit row.

In the first round, the head did not beat a hand-written rule table, and the difference sat within noise. After retraining, the head scored 0.728 on fresh cases. The rule table reached 0.566 on the same cases. The lead was 0.161. Its interval runs from 0.131 to 0.191.

The retrained head closed cases wrongly at a rate of 0.047, above the bar set in advance. The cause was in the evidence pack the test environment built. A fix came only after the bar was missed. With it the rate was 0.0113, and both readings are published.

Jev, a hosted service, answered a single question in 0.35 seconds. Under four parallel streams, the open head's median response was 3.1 seconds.

On a second business process, the head's lead was gone. The code is public under the Apache 2.0 licence. The finding covers one decision task on two synthetic processes built the same way.

10. The Delegation Index

Five of the campaigns feed one quarterly number, the Delegation Index. Its scale runs from zero to a hundred. A reading of fifty means the cost of handing work to agents equals what oversight returns at a fixed budget. The first reading is 47.2. The band runs from 40.8 to 53.2.

The reading rests on 20,154 units behind the five numbers, of different kinds. The sprawl component scores lowest, at 26.7. The band contains fifty, so this quarter's evidence fits balance and imbalance alike. The index covers the lab's own synthetic measurements, and market and firm data lie outside its scope.

Ask for the registered prediction, with its date. Which band did each number have to land in? Someone must name the result that would have stopped the work. The figures need an independent recompute from the raw files, using fresh code. Request the negative results as well.

The campaigns and their corrections are at dinand.com slash research.