What it means
An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds a gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. A Trial Balance applies the same idea to agents.
The layer that hands out the work writes down what the agent owes. The harness writes down what the agent actually ran. The auditor reads only where the two books fail to balance, and then reads only the requisition that failed.
In the campaign, a procurement workflow ran at lengths from one requisition to sixteen, with faults injected into each diary. At eight requisitions, the balance check reached 0.80 omission detection at 745 prompt tokens, with a 90% interval of 740 to 749. Plain log review never reached 0.80 at any budget or length, and its best was 0.58. Per thousand tokens read, the ledger found about 21 times as many correct omissions.
The ledger has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, the balance is clean, and the check found 0 of 1,044 of them. A full read of the log finds those almost every time. Each audit covers faults the other misses, so a deployment that needs both needs both.