SIGIL: does a liability ledger change what an overseer authorizes?
Refuted30,000 supervised episodes
The registered prediction that review pays on irreversible decisions failed in all 12 series; a liability ledger raised wrongful blocking at the top tier to 42.7%.
Registration and verificationOutput contract frozen before any harness code existed; both cost models run in full. Recomputed independently from the raw episode ledgers with zero mismatches. Published as a refutation.
The question
If a reviewer is going to hold the power to stop an AI worker, the usual next step is to make it accountable for what it lets through. This campaign asked whether an AI supervisor authorises differently when it carries a running record of the damage its own past approvals caused, and how the value of its review changes as decisions become harder to undo.
What was registered
The hypothesis, H5, came from a working paper on accountability. It predicted that review before commitment loses value on reversible decisions, turns positive past some level of irreversibility and rises steeply from there. It is falsified by a flat response, a concave one or the absence of any crossing.
The output contract was frozen on 30 August, before any harness code existed. It fixed the scoring, a gate requiring the unsupervised worker to err between 10% and 35% on a pilot, and a mandatory rerun under a second, linear cost model. One amendment, recorded before the second pilot, changed only the task generator and pre-committed a stopping rule against retuning more than once.
What was measured
An AI worker approved or rejected synthetic firm decisions across five tiers, from invoice approvals that reverse next cycle to public statements that cannot be taken back. The same base model, in a second seat, reviewed it under four designs:
no oversight;
an after-the-fact audit of a 20% sample;
review with the power to block before commitment;
the same blocking review carrying a liability ledger of the harm its own past approvals caused.
Worker and reviewer share one model on purpose, so any difference comes from the incentive design. Every decision was scored against programmatic ground truth. 30,000 episodes ran across two full campaigns, one per cost model, with three seeds per cell.
What it found
Review never crossed zero, in any of the 12 registered series, under either cost model. Its value collapsed at the most irreversible tier, the place the theory wanted it to rise: −2.64 per decision against no oversight. The liability ledger made it worse, at −3.81. At that tier plain review wrongly blocked 30.1% of decisions and the ledger pushed it to 42.7%. Refusal was the one action the ledger never charged for.
The mechanism is measurable. A reviewer that can only block can fix only one kind of mistake, a wrongful approval. The share of worker errors that were wrongful approvals fell from 93% at tier 1 to 1.4% at tier 5 (the executive piece rounds this to 2%; the paper gives 1 of 69), so at the decisions where buyers most want review there was almost nothing for a blocker to catch. Every table was recomputed independently from the raw episode ledgers, with zero mismatches on 14,950 convex and 14,952 linear scored units (29,902 in all).
Headline numbers, as published
Registered series showing the predicted zero crossing in the value of review
0 of 12
Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model)
−3.81 against −2.64
Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger
30.1% against 42.7%
Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 5
93% to 1.4%
Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived)
A controlled simulation with one open-weight model family in both seats, scored against constructed ground truth.
The liability ledger is a thin, simulated form of accountability.
A follow-up registered to test stronger models stopped without running an oversight trial: four other models got between 97.3% and 100% of cases right unsupervised, which left a reviewer nothing to catch. Read the result as specific to this setting.