Methods
The protocol
The lab studies how organisations delegate work to AI agents and how they oversee it. Every campaign follows the same eight steps, and the steps are built so the lab cannot quietly steer a result toward the answer it hoped for.
-
Ask one question
A campaign starts with a question a buyer or a regulator would recognise, such as whether a second model checks the first or what a liability ledger does to a reviewer. It names the quantity that would answer it and the null result that would mean the idea adds nothing.
-
Register the prediction, the bands and the kill conditions
Before any unit runs, the lab writes down the hypotheses, the band each number has to land in to count as a hit, the statistics that will be used and the conditions that stop the run. The document is dated and frozen. A later change goes in as a dated amendment that says what moved and why, and it is written before the data it affects is read.
Each campaign tests its instrument first. A small calibration stage checks that the task is neither trivial nor impossible and that the harness records what it should. That gate is allowed to test whether the instrument works. It is never allowed to require the result the campaign exists to measure; the lab learned that by halting a campaign on exactly that mistake.
-
Attack the design
A reviewer whose job is to break the headline goes over the design and the analysis: a confound, a statistic that hides a split, a ceiling or a floor, a scoring rule that favours one arm. What the review finds is fixed before the run or stated as a limit in the piece.
-
Run to the registered rule
The campaign runs on open models on hardware the lab owns, so any run can be repeated. It stops where the registration says it stops. A failed gate, an unmet validity check or a spent compute budget ends the run, and what is complete at that point is analysed as it stands. The lab does not extend a run, loosen a threshold or add a round to rescue it.
-
Verify from the raw files
A separate pass, with fresh code and no access to the original analysis, recomputes the reported numbers from the raw records. Each piece says how many numbers were checked and how many matched. Where the two passes disagree, the disagreement is chased down and the piece carries the verified value. On the sprawl-cost campaign this pass found duplicate runs and moved a verdict from supported to refuted before publication.
-
Publish the negative results
A refuted hypothesis is published with the same care as a confirmed one, next to the registered prediction it contradicted. A null result is reported as a bound, with the size of effect the study could have detected. What was measured is kept apart from what the lab infers, and an explanation is labelled as one until a test has been run against it.
-
Release the code
Each campaign keeps its registration, its harness, its unmodified ledger and its analysis together, so the published numbers can be traced to the rows that produced them. The code and records are released with the pieces as each campaign is cleared for it.
-
Log every correction
When a published number turns out to be wrong, the correction is made in the open and dated, and the original number stays visible beside it. The measured claims index keeps a row for every headline number, so a correction has a fixed place to land.