Build · from the Trial Balance campaign
The Delegation Ledger
Before an agent starts, the dispatcher notes each step it owes. This page sets that list against the steps the tool runner saw execute and shows you only where they part.
Simulation from published coefficients, not a deployed system
sample: demo/sample-ledger.jsonl, written by demo/make_sample.py; amounts and times are made up for the demo
Where the books disagree
Trial balance
one line per task, agent and step
the tables scroll sideways
| task | agent | step | owed | ran | owed amount | ran amount | status | reason | due |
|---|
By agent
each agent is its own account
| agent | tasks | lines | owed | ran | ok | open | skipped | amount | unowed |
|---|
By task
tasks with a line out of balance first
| task | agents | lines | ok | open | skipped | amount | unowed | closed | first line | last line |
|---|
Timeline per task
owe and ran lines in time order
square above the rule: owed · circle below: ran · bar across: task closed · a step out of balance: heavy diamond above, large ringed circle below · each mark has a title with its line
What each column means
read this before the numbers
What the generator injected
sample-truth.json
The sample carries one requisition per fault condition from the campaign, plus one in progress and one the commissioner closed early. Each verdict says whether the balance above shows that fault.
| task | condition | detail | in the balance |
|---|
What this cannot see
Blind spot by design
A wrong amount carried on both sides balances. If the error got in before the work was assigned, the obligation and the run agree on the wrong value, and the line reads clean. Requisition R-0107 in the sample is one of these: its quote was 6,418 and both books carry 4,618.
In the campaign the balance check found 0 of 1,044 such faults. A full read of the log found them almost every time, because the vendor's quote still showed the true amount.
source: Dinand Tinholt, Research, "A Trial Balance for Agent Omissions", campaign page, found section, published 2026-10-01What the campaign measured
18,360 audits, one auditor model
| 0.80 | omission detection reached by reading the imbalance first, at 745 prompt tokens and eight requisitionsclaims.yaml tb-c80; 90% interval 740 to 749 tokens |
| 0.58 | best omission detection by plain log review at any budget or trace length; it never reached 0.80claims.yaml tb-lr-best |
| 21× | correct omission findings per thousand tokens read, imbalance-first against log review at its best budgetclaims.yaml tb-ratio; 1.34 against 0.063, eight requisitions |
| 0 / 1,044 | wrong amounts carried on both sides that the balance check foundcampaign page, trial-balance, found section |
These figures come from the campaign's audits of a seven-step procurement workflow. The library on this page keeps the same two books and computes the same balance, and nobody has yet measured it on a live system.