Measured claims

Every public headline number

One row per number the lab has published, newest first, with the interval where the piece gives one. Quote the number with its campaign and link to the piece. 60 claims.

DateClaimNumberIntervalCampaignSource
2026-10-07First round: open head decision-class accuracy, 800 eval cases, zero-shot0.545noneThe open decision headExecutive piece
2026-10-07First round: rule table decision-class accuracy, 800 eval cases0.575noneThe open decision headExecutive piece
2026-10-07First round: open head minus rule table, 800 eval cases (inconclusive)-0.029+0.001 to -0.060The open decision headExecutive piece
2026-10-07Jev (hosted) decision-class accuracy on the same 800 eval cases0.540noneThe open decision headExecutive piece
2026-10-07Jev decision-class accuracy on the 200-case contradictory-evidence set0.536noneThe open decision headExecutive piece
2026-10-07Second round: retrained head decision-class accuracy, 800 fresh cases (eval2)0.728noneThe open decision headExecutive piece
2026-10-07Second round: rule table decision-class accuracy, 800 fresh cases (eval2)0.566noneThe open decision headExecutive piece
2026-10-07Second round: retrained head decision-class accuracy, harder set0.739noneThe open decision headExecutive piece
2026-10-07Second round: rule table decision-class accuracy, harder set0.565noneThe open decision headExecutive piece
2026-10-07Retrained head ahead of Jev's 0.536 on the 199 cases both scored, harder set (gap excludes zero)0.199noneThe open decision headExecutive piece
2026-10-07Smaller fine-tuned head decision-class accuracy, fresh set (experimental label)0.701noneThe open decision headExecutive piece
2026-10-07Wrong-claim rate, retrained head, fresh set, before the evidence-pack fix (above the registered trigger)0.047noneThe open decision headExecutive piece
2026-10-07Wrong-claim rate, retrained head, after the evidence-pack fix0.0113noneThe open decision headExecutive piece
2026-10-07Wrong-claim rate, rule table, after the evidence-pack fix0.0088noneThe open decision headExecutive piece
2026-10-07Wrong-claim rate, learned head, after the evidence-pack fix0.0163noneThe open decision headExecutive piece
2026-10-07Expected calibration error of the first head (registered bar 0.08)about 0.52noneThe open decision headExecutive piece
2026-10-07Item-level calibration error of the later judging head, full eval2: uncalibrated 0.0178, refit0.0143noneThe open decision headExecutive piece
2026-10-07Jev single-question response time at $0.0000128 for 304 input tokens, measured 24 September 20260.35 secondsnoneThe open decision headExecutive piece
2026-10-07Open head fast-mode median response time, four benchmark streams3.1 secondsnoneThe open decision headExecutive piece
2026-10-07Open head fast-mode p90 response time, four benchmark streams4.8 secondsnoneThe open decision headExecutive piece
2026-10-07Checkpoints logged in the J5 eval2 run, across 794 cases (about 3.4 per case)2,681noneThe open decision headExecutive piece
2026-10-07Share of repeats that disagree, Jev, first run (open head 4.95% eval, 5.85% challenge, with a second job on the endpoint)0.05%noneThe open decision headExecutive piece
2026-10-07Share of repeats that disagree, open head alone at one call in flight (both processes)0.05%noneThe open decision headExecutive piece
2026-10-07Share of repeats that disagree, open head at four calls in flight, onboarding (accounts payable 1.95%)3.4%noneThe open decision headExecutive piece
2026-10-07Second process (accounts payable): open head accuracy on ap-eval, against rule table 0.6480.650noneThe open decision headExecutive piece
2026-10-07Second process: learned head trained on the first process, no retraining, ap-eval0.522noneThe open decision headExecutive piece
2026-10-07Second process: learned head retrained on 100 of the new process's labelled cases (271 decisions), ap-eval0.663noneThe open decision headExecutive piece
2026-10-07Laya fine-tuned on 1,567 labelled decisions, onboarding harder set (rule table 0.565)0.815noneThe open decision headExecutive piece
2026-10-07Laya fine-tuned on 201 labelled decisions from 58 accounts-payable cases, ap-challenge, below the rule table0.554noneThe open decision headExecutive piece
2026-10-07AUC of a vote across several readings as a wrong-claim guard (not in use)0.56noneThe open decision headExecutive piece
2026-10-07AUC of a guard that read the cited document itself (not in use)0.52noneThe open decision headExecutive piece
2026-10-01Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions745740 to 749A Trial Balance for Agent OmissionsPaper
2026-10-01Best omission detection by plain log review at any budget or trace length (0.80 never reached)0.58noneA Trial Balance for Agent OmissionsPaper
2026-10-01Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitionsabout 21 times (1.34 against 0.063)noneA Trial Balance for Agent OmissionsPaper
2026-10-01Omission detection by plain log review reading the whole diary0.250.19 to 0.31A Trial Balance for Agent OmissionsPaper
2026-10-01Omission detection reading the whole diary with a per-step checklist instruction0.580.48 to 0.67A Trial Balance for Agent OmissionsPaper
2026-10-01Wrong amounts carried on both sides of the book found by the balance check0 of 1,044noneA Trial Balance for Agent OmissionsPaper
2026-10-01Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179)0.7980.738 to 0.857A Trial Balance for Agent OmissionsPaper
2026-10-01Omission detection, imbalance-first over a random sample of the same reading budget+0.9070.891 to 0.925A Trial Balance for Agent OmissionsPaper
2026-09-17Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust)1.46651.4101 to 1.5230Sprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-17Sprawl cost exponent in fleet size, per-agent monitoring with review that scales0.90010.8664 to 0.9337Sprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-17Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring2.76× against 1.87×noneSprawl Cost Under a Fixed Monitoring BudgetExecutive piece
2026-09-17Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical55 of 64noneSprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-17Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents2.78 to 22.49noneSprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-17Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145)1.4949 against 1.32381.4037 to 1.5862; 1.2294 to 1.4182Sprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-17Banked runs the independent verification found to be duplicates95 duplicate pairs (190 of 256 rows; 161 distinct simulations)noneSprawl Cost Under a Fixed Monitoring BudgetPaper
2026-09-10Registered series showing the predicted zero crossing in the value of review0 of 12noneSIGIL: does a liability ledger change what an overseer authorizes?Paper
2026-09-10Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model)−3.81 against −2.64noneSIGIL: does a liability ledger change what an overseer authorizes?Paper
2026-09-10Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger30.1% against 42.7%noneSIGIL: does a liability ledger change what an overseer authorizes?Paper
2026-09-10Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 593% to 1.4%noneSIGIL: does a liability ledger change what an overseer authorizes?Paper
2026-09-10Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived)0 of 14,950 convex and 0 of 14,952 linearnoneSIGIL: does a liability ledger change what an overseer authorizes?Paper
2026-09-03Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948)−1.50noneDelegation CliffPaper
2026-09-03Change in quality margin per increment of reviewer capacity, the largest positive lever+1.46noneDelegation CliffPaper
2026-09-03Model capability against agent self-check calibration (read as tied)+0.95 against +0.92noneDelegation CliffPaper
2026-09-03Share of the design space in a broad transition band at a 0.95 quality floor (no cliff)60%noneDelegation CliffPaper
2026-09-03Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep)79%noneDelegation CliffPaper
2026-09-03Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.050.0117noneDelegation CliffPaper
2026-09-02Error correlation (phi), identical weights on different silicon, contested items only+0.447noneError Independence: does different hardware buy a second opinion?Paper
2026-09-02Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280)−0.102noneError Independence: does different hardware buy a second opinion?Paper
2026-09-02Error correlation (phi) across all pairs before conditioning on item difficulty0.556 to 0.827noneError Independence: does different hardware buy a second opinion?Paper

Intervals are 95% unless the campaign page says otherwise; Trial Balance reports 90% cluster-bootstrap intervals. A number that a later correction changes keeps its row, and the correction is added beside it.

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top