Instruments · SIGIL: does a liability ledger change what an overseer authorizes?
SIGIL tier explorer
Pick an irreversibility tier and compare what blocking review and review with a liability ledger did to the value of oversight there, beside what the registered prediction said would happen.
What it computes
For each of five irreversibility tiers, from invoices that can be undone next cycle to public corrections that cannot be unsaid, the net value per decision of two oversight designs against no oversight at all, the share of decisions each design blocked, the share it blocked wrongly, and the share of worker errors a blocker could have caught.
The formula in plain words
Net value per decision is the negative of loss plus oversight cost, averaged over the tier, minus the same average with no oversight. Under the convex cost model an uncaught wrongful approval at tier t costs t squared, a wrongful rejection or a wrongful block costs 35% of that, and every consultation of the supervisor costs 0.15. Values are pooled over three seeds. Below zero, oversight costs more than it saves.
What this does not say
It does not say that review never pays. The result is local to one base model, qwen3.5, sitting in both seats, on synthetic checklist-compliance tasks, with three seeds, and the liability ledger is a thin, simulated form of accountability. Tier 3 produced no worker errors at all, so its value is the consultation fee and nothing more. The post-hoc audit design, which trimmed the worst losses, is outside this view.
Source
Every number comes from Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers (Paper, Medium, 2026-09-10), with some from Skin in the Game Made the AI Supervisor Worse (Executive piece, 2026-09-09). The campaign record is SIGIL: does a liability ledger change what an overseer authorizes?.
Every number this instrument uses
| Number | As published | Where |
|---|---|---|
| Task family at each tier | Invoice approval, reversible next cycle / Refund release, recoverable with effort / Contract clause acceptance, binds until renegotiated / Data deletion, recoverable only from backup if at all / Public correction, cannot be unsaid | Paper, Section 3.1, The delegation game |
| Net value per decision against no oversight, tiers 1 to 5 | review: −0.13, −0.17, −0.18, −0.17, −2.64; ledger: −0.14, −0.20, −0.16, −0.51, −3.81 | Executive piece, The one table to keep (convex cost model, pooled over seeds) |
| Share of decisions blocked and wrongly blocked, tiers 1 to 5, in percent | review block: 7.2, 21.6, 2.5, 3.5, 30.5; review wrong: 3.3, 11.9, 1.1, 1.5, 30.1; ledger block: 3.9, 4.3, 1.6, 6.0, 43.4; ledger wrong: 1.5, 2.5, 0.3, 5.9, 42.7 | Paper, Table 2, supervisor blocking per tier, convex campaign |
| Worker errors a blocker could catch (wrongful approvals), tiers 1 to 5; tier 3 had no errors | share: 93, 56, , 17, 1.4; approves: 157, 115, 0, 17, 1; rejects: 12, 89, 0, 81, 68 | Paper, Figure 2, false-approve share of autonomous worker errors, convex campaign |
| Registered series that showed the predicted zero crossing | 0 of 12 | Paper, Abstract and section 4.1 |
| Second difference at tier 4, review and ledger (the prediction needed it positive) | review: −2.49; ledger: −2.95 | Paper, Section 4.1 |
| Tier-5 value under the linear cost model, review and ledger, pooled | review: −0.656; ledger: −0.889 | Paper, Section 4.1, linear campaign |
| Tier 5: review's mean net value against an always-allow replay of the same episodes | actual: −3.47; allow: −0.91 | Paper, Section 4.2 |
| Registered band for the unsupervised worker's error rate, and the rate the campaigns ran at (convex, linear), in percent | low: 10; high: 35; convex: 14.4; linear: 14.5 | Paper, Section 2 and section 3.5, the smoke gate |
| What the registered prediction said at each tier, and how it fared | tier1: Predicted: review loses money on reversible decisions. Held: review sits at −0.13 and the ledger at −0.14, and every seed shows the loss. It is the only clause of the prediction the data grants.; tier2: Predicted: value climbs toward a zero crossing somewhere along the gradient. Refuted: both designs stay below zero, and the ledger loses more than plain review.; tier3: Predicted: value climbs toward a zero crossing. Refuted, and this tier cannot speak to it: the worker made zero errors here across 6, 000 episodes, so what remains is the consultation fee.; tier4: Predicted: past the crossing, value rises convexly. Refuted: no series crossed zero, and the second difference at tier 4 is −2.49 for review and −2.95 for the ledger, curving down.; tier5: Predicted: the largest gain of all, at the most irreversible decisions. Refuted: value collapses to −2.64 for review and −3.81 for the ledger, the same sign in every seed and under both cost models. | Paper, Section 1 and section 4.1, hypothesis H5 |