Measured claims
Every public headline number
One row per number the lab has published, newest first, with the interval where the piece gives one. Quote the number with its campaign and link to the piece. 60 claims.
| Date | Claim | Number | Interval | Campaign | Source |
|---|---|---|---|---|---|
| 2026-10-07 | First round: open head decision-class accuracy, 800 eval cases, zero-shot | 0.545 | none | The open decision head | Executive piece |
| 2026-10-07 | First round: rule table decision-class accuracy, 800 eval cases | 0.575 | none | The open decision head | Executive piece |
| 2026-10-07 | First round: open head minus rule table, 800 eval cases (inconclusive) | -0.029 | +0.001 to -0.060 | The open decision head | Executive piece |
| 2026-10-07 | Jev (hosted) decision-class accuracy on the same 800 eval cases | 0.540 | none | The open decision head | Executive piece |
| 2026-10-07 | Jev decision-class accuracy on the 200-case contradictory-evidence set | 0.536 | none | The open decision head | Executive piece |
| 2026-10-07 | Second round: retrained head decision-class accuracy, 800 fresh cases (eval2) | 0.728 | none | The open decision head | Executive piece |
| 2026-10-07 | Second round: rule table decision-class accuracy, 800 fresh cases (eval2) | 0.566 | none | The open decision head | Executive piece |
| 2026-10-07 | Second round: retrained head decision-class accuracy, harder set | 0.739 | none | The open decision head | Executive piece |
| 2026-10-07 | Second round: rule table decision-class accuracy, harder set | 0.565 | none | The open decision head | Executive piece |
| 2026-10-07 | Retrained head ahead of Jev's 0.536 on the 199 cases both scored, harder set (gap excludes zero) | 0.199 | none | The open decision head | Executive piece |
| 2026-10-07 | Smaller fine-tuned head decision-class accuracy, fresh set (experimental label) | 0.701 | none | The open decision head | Executive piece |
| 2026-10-07 | Wrong-claim rate, retrained head, fresh set, before the evidence-pack fix (above the registered trigger) | 0.047 | none | The open decision head | Executive piece |
| 2026-10-07 | Wrong-claim rate, retrained head, after the evidence-pack fix | 0.0113 | none | The open decision head | Executive piece |
| 2026-10-07 | Wrong-claim rate, rule table, after the evidence-pack fix | 0.0088 | none | The open decision head | Executive piece |
| 2026-10-07 | Wrong-claim rate, learned head, after the evidence-pack fix | 0.0163 | none | The open decision head | Executive piece |
| 2026-10-07 | Expected calibration error of the first head (registered bar 0.08) | about 0.52 | none | The open decision head | Executive piece |
| 2026-10-07 | Item-level calibration error of the later judging head, full eval2: uncalibrated 0.0178, refit | 0.0143 | none | The open decision head | Executive piece |
| 2026-10-07 | Jev single-question response time at $0.0000128 for 304 input tokens, measured 24 September 2026 | 0.35 seconds | none | The open decision head | Executive piece |
| 2026-10-07 | Open head fast-mode median response time, four benchmark streams | 3.1 seconds | none | The open decision head | Executive piece |
| 2026-10-07 | Open head fast-mode p90 response time, four benchmark streams | 4.8 seconds | none | The open decision head | Executive piece |
| 2026-10-07 | Checkpoints logged in the J5 eval2 run, across 794 cases (about 3.4 per case) | 2,681 | none | The open decision head | Executive piece |
| 2026-10-07 | Share of repeats that disagree, Jev, first run (open head 4.95% eval, 5.85% challenge, with a second job on the endpoint) | 0.05% | none | The open decision head | Executive piece |
| 2026-10-07 | Share of repeats that disagree, open head alone at one call in flight (both processes) | 0.05% | none | The open decision head | Executive piece |
| 2026-10-07 | Share of repeats that disagree, open head at four calls in flight, onboarding (accounts payable 1.95%) | 3.4% | none | The open decision head | Executive piece |
| 2026-10-07 | Second process (accounts payable): open head accuracy on ap-eval, against rule table 0.648 | 0.650 | none | The open decision head | Executive piece |
| 2026-10-07 | Second process: learned head trained on the first process, no retraining, ap-eval | 0.522 | none | The open decision head | Executive piece |
| 2026-10-07 | Second process: learned head retrained on 100 of the new process's labelled cases (271 decisions), ap-eval | 0.663 | none | The open decision head | Executive piece |
| 2026-10-07 | Laya fine-tuned on 1,567 labelled decisions, onboarding harder set (rule table 0.565) | 0.815 | none | The open decision head | Executive piece |
| 2026-10-07 | Laya fine-tuned on 201 labelled decisions from 58 accounts-payable cases, ap-challenge, below the rule table | 0.554 | none | The open decision head | Executive piece |
| 2026-10-07 | AUC of a vote across several readings as a wrong-claim guard (not in use) | 0.56 | none | The open decision head | Executive piece |
| 2026-10-07 | AUC of a guard that read the cited document itself (not in use) | 0.52 | none | The open decision head | Executive piece |
| 2026-10-01 | Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions | 745 | 740 to 749 | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Best omission detection by plain log review at any budget or trace length (0.80 never reached) | 0.58 | none | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitions | about 21 times (1.34 against 0.063) | none | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Omission detection by plain log review reading the whole diary | 0.25 | 0.19 to 0.31 | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Omission detection reading the whole diary with a per-step checklist instruction | 0.58 | 0.48 to 0.67 | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Wrong amounts carried on both sides of the book found by the balance check | 0 of 1,044 | none | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179) | 0.798 | 0.738 to 0.857 | A Trial Balance for Agent Omissions | Paper |
| 2026-10-01 | Omission detection, imbalance-first over a random sample of the same reading budget | +0.907 | 0.891 to 0.925 | A Trial Balance for Agent Omissions | Paper |
| 2026-09-17 | Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust) | 1.4665 | 1.4101 to 1.5230 | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-17 | Sprawl cost exponent in fleet size, per-agent monitoring with review that scales | 0.9001 | 0.8664 to 0.9337 | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-17 | Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring | 2.76× against 1.87× | none | Sprawl Cost Under a Fixed Monitoring Budget | Executive piece |
| 2026-09-17 | Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical | 55 of 64 | none | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-17 | Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents | 2.78 to 22.49 | none | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-17 | Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145) | 1.4949 against 1.3238 | 1.4037 to 1.5862; 1.2294 to 1.4182 | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-17 | Banked runs the independent verification found to be duplicates | 95 duplicate pairs (190 of 256 rows; 161 distinct simulations) | none | Sprawl Cost Under a Fixed Monitoring Budget | Paper |
| 2026-09-10 | Registered series showing the predicted zero crossing in the value of review | 0 of 12 | none | SIGIL: does a liability ledger change what an overseer authorizes? | Paper |
| 2026-09-10 | Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model) | −3.81 against −2.64 | none | SIGIL: does a liability ledger change what an overseer authorizes? | Paper |
| 2026-09-10 | Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger | 30.1% against 42.7% | none | SIGIL: does a liability ledger change what an overseer authorizes? | Paper |
| 2026-09-10 | Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 5 | 93% to 1.4% | none | SIGIL: does a liability ledger change what an overseer authorizes? | Paper |
| 2026-09-10 | Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived) | 0 of 14,950 convex and 0 of 14,952 linear | none | SIGIL: does a liability ledger change what an overseer authorizes? | Paper |
| 2026-09-03 | Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948) | −1.50 | none | Delegation Cliff | Paper |
| 2026-09-03 | Change in quality margin per increment of reviewer capacity, the largest positive lever | +1.46 | none | Delegation Cliff | Paper |
| 2026-09-03 | Model capability against agent self-check calibration (read as tied) | +0.95 against +0.92 | none | Delegation Cliff | Paper |
| 2026-09-03 | Share of the design space in a broad transition band at a 0.95 quality floor (no cliff) | 60% | none | Delegation Cliff | Paper |
| 2026-09-03 | Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep) | 79% | none | Delegation Cliff | Paper |
| 2026-09-03 | Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.05 | 0.0117 | none | Delegation Cliff | Paper |
| 2026-09-02 | Error correlation (phi), identical weights on different silicon, contested items only | +0.447 | none | Error Independence: does different hardware buy a second opinion? | Paper |
| 2026-09-02 | Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280) | −0.102 | none | Error Independence: does different hardware buy a second opinion? | Paper |
| 2026-09-02 | Error correlation (phi) across all pairs before conditioning on item difficulty | 0.556 to 0.827 | none | Error Independence: does different hardware buy a second opinion? | Paper |
Intervals are 95% unless the campaign page says otherwise; Trial Balance reports 90% cluster-bootstrap intervals. A number that a later correction changes keeps its row, and the correction is added beside it.