Dinand Tinholt · Experiments

Builds · open library and live view

Build · from the Trial Balance campaign

The Delegation Ledger

Before an agent starts, the dispatcher notes each step it owes. This page sets that list against the steps the tool runner saw execute and shows you only where they part.

Simulation from published coefficients, not a deployed system

sample: demo/sample-ledger.jsonl, written by demo/make_sample.py; amounts and times are made up for the demo

Drop a ledger file here, or several segments of a rotated one.

Paste a ledger instead
loading ledger

Where the books disagree

Trial balance

one line per task, agent and step

the tables scroll sideways

Trial balance
taskagentstepowedranowed amountran amountstatusreasondue

By agent

each agent is its own account

Lines per agent by status
agenttaskslinesowedranokopenskippedamountunowed

By task

tasks with a line out of balance first

Lines per task by status
taskagentslinesokopenskippedamountunowedclosedfirst linelast line

Timeline per task

owe and ran lines in time order

square above the rule: owed · circle below: ran · bar across: task closed · a step out of balance: heavy diamond above, large ringed circle below · each mark has a title with its line

What each column means

read this before the numbers

What the generator injected

sample-truth.json

The sample carries one requisition per fault condition from the campaign, plus one in progress and one the commissioner closed early. Each verdict says whether the balance above shows that fault.

Injected faults and whether the balance shows them
taskconditiondetailin the balance

What this cannot see

Blind spot by design

A wrong amount carried on both sides balances. If the error got in before the work was assigned, the obligation and the run agree on the wrong value, and the line reads clean. Requisition R-0107 in the sample is one of these: its quote was 6,418 and both books carry 4,618.

In the campaign the balance check found 0 of 1,044 such faults. A full read of the log found them almost every time, because the vendor's quote still showed the true amount.

source: Dinand Tinholt, Research, "A Trial Balance for Agent Omissions", campaign page, found section, published 2026-10-01

What the campaign measured

18,360 audits, one auditor model

0.80omission detection reached by reading the imbalance first, at 745 prompt tokens and eight requisitionsclaims.yaml tb-c80; 90% interval 740 to 749 tokens
0.58best omission detection by plain log review at any budget or trace length; it never reached 0.80claims.yaml tb-lr-best
21×correct omission findings per thousand tokens read, imbalance-first against log review at its best budgetclaims.yaml tb-ratio; 1.34 against 0.063, eight requisitions
0 / 1,044wrong amounts carried on both sides that the balance check foundcampaign page, trial-balance, found section

These figures come from the campaign's audits of a seven-step procurement workflow. The library on this page keeps the same two books and computes the same balance, and nobody has yet measured it on a live system.