Reading where a commitment ledger fails to balance found 0.80 of skipped steps at 745 tokens of reading; log review never reached 0.80 at any budget.
Registration and verificationPre-registered before the first production unit, after a first attempt halted at its calibration gate. Independently rescored from the raw ledgers: point estimates equal to the registered analysis and the verifier's, interval bounds differing only by bootstrap noise. One auditor model, stated as a limit.
The question
An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds the gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. This campaign asked whether the same idea works for agents: if the layer that hands out the work writes down what the agent owes, and the harness writes down what the agent actually ran, can an auditor find skipped steps by reading only where the two books fail to balance, and do it for a fraction of what reading the log costs?
What was registered
The registration said imbalance-first audit would reach 0.80 omission detection at a third of log review's reading cost or less, at 8, 12 and 16 requisitions. It also predicted that a balanced book hides a wrong amount carried on both sides, and that a reader of the log would miss a skipped step far more often when the agent wrote a line claiming it was done.
The first attempt halted at its calibration gate on 11 September with zero production units. Its gate had required a reading to move that the campaign existed to measure, so a strong true result would have failed it. The lab rewrote its rule the same evening: a gate tests only whether the instrument works. The second attempt was registered and frozen before its first unit, with one added arm, a log reader told to tick off every step the procedure requires before it gives a verdict.
What was measured
A procurement workflow with seven steps ran at seven lengths, from one requisition to sixteen. Five conditions were injected into each diary: no fault, a skipped step, a skipped step with a line claiming it was done, a wrong amount, and a wrong amount carried on both sides of the book. Four auditors read each diary under seven reading budgets from 200 to 12,800 tokens.
Log review reads the diary from the top.
Imbalance-first reads the trial-balance report and then only the requisition that failed to balance.
A random sample of the same size separates reading less from reading the right thing.
The checklist reader gets the same diary as log review plus the instruction to mark every owed step present or absent.
The auditor was qwen3.5 under a fixed reply schema: 18,360 audits, including 3,240 on diaries written by a different model running the workflow for real. An independent rescoring from the raw ledgers found point estimates equal to the registered analysis and the verifier's, with interval bounds differing only by bootstrap noise.
What it found
Imbalance-first reached 0.80 omission detection at the smallest budget, 745 prompt tokens at eight requisitions, and sat between 0.92 and 1.00 in every cell. Log review never reached 0.80 at any budget or length; its best was 0.58. Per thousand tokens read, imbalance-first produced about 21 times as many correct omission findings as log review at its best budget.
The balance check has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, the trial balance is clean, and the check found 0 of 1,044 of them. A full read of the diary finds those almost every time, because the vendor's quote still shows the true amount. The two audits cover different faults, and a deployment that needs both needs both.
The checklist instruction is the cheapest result here. With the whole diary in front of it, plain log review found skipped steps at 0.25; the same auditor told to tick off each owed step found 0.58, and accused no clean diary.
Headline numbers, as published
Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions
745[740 to 749]
Best omission detection by plain log review at any budget or trace length (0.80 never reached)
0.58
Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitions
about 21 times (1.34 against 0.063)
Omission detection by plain log review reading the whole diary
0.25[0.19 to 0.31]
Omission detection reading the whole diary with a per-step checklist instruction
0.58[0.48 to 0.67]
Wrong amounts carried on both sides of the book found by the balance check
0 of 1,044
Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179)
0.798[0.738 to 0.857]
Omission detection, imbalance-first over a random sample of the same reading budget
+0.907[0.891 to 0.925]
Share of injected faults named correctly at a full read, by fault class, synthetic primary seeds, 84 per class. Rebuilt from the numbers in A Trial Balance for Agent Omissions.Show the numbers as a table
Fault class
Imbalance-first
Checklist reader
Log review
Silent omission
1.00
0.73
0.48
Phantom completion
1.00
0.36
0.18
Value commission
1.00
0.98
0.98
Consistent commission
0.00
0.95
1.00
What it does not claim
One auditor model. The second registered family was excluded before its calibration gate because the host could not supply its power draw, and nothing about it is claimed.
The registered cost ratio is unbounded because log review never reached the bar. The 21-times figure is the finite comparison to quote.
The campaign cannot say how much of the checklist's gain comes from the list and how much from the sentence defining what counts as evidence of a step.
One workflow, one fault per diary, and real diaries only up to three requisitions.
The cost of building a commissioning layer that writes obligations in ledger form was not priced.