Ideas across time · Research

Research

Simulations I run when I want to know if an idea about AI survives contact with real numbers.

Pre-registered campaigns

The questions I put to the test

Each campaign writes down its predictions before a single unit runs, then reports what held, what failed and what the result does not claim. 7 campaigns so far.

Annual report · 2026

The Measured Enterprise 2026

The lab’s annual report: the Delegation Index, the protocol, seven pre-registered campaigns, the corrections logged and every measured claim, in one document.

Read the report

Three-part publication · September 2026

The Decision Head

A self-hosted service that answers a workflow's typed questions with a calibrated probability. The executive briefing, the full paper and the blog post that started it.

Open the publication

Mixed

The open decision head

A retrained open decision head scored 0.728 against 0.566 for a rule table on 800 fresh cases, a gap that excludes zero; the first round missed, with 0.545 against 0.575 (gap -0.029, interval +0.001 to -0.060).

Read the campaign

Lab notebook

Studies & simulations

16 published studies, most of them pre-registered simulations: agent oversight, delegation, supply-chain stress and multi-agent markets.

Medium ↗

A Trial Balance for Agent Omissions

Agent logs record what happened and have no row for what didn’t. A pre-registered test of commitment accounting, an obligation ledger reconciled against executed calls, for catching omissions in procurement-workflow traces.

Oversight & accountability

Medium ↗

Pricing Agent Autonomy

A simulation sweeping twelve dimensions of organisational design to price the bet that moving work further from human review is cheaper. Delegation depth carries the strongest weight, and it is negative.

Oversight & accountability

Medium ↗

When the Second Opinion Shares the Blind Spot

One model on two machines, with different quantisation and serving stacks, made the same mistakes. What separates one reviewer from another is the model family: a case for dissimilar redundancy.

Oversight & accountability

Medium ↗

A ledger for agentic decisions

Difficulty is a property of the work; risk is a property of the consequences. Testing, in simulation and as a product, the rule that decides which agentic decisions a person must see.

Oversight & accountability

Medium ↗

The Breaking Point

45,312 simulated supply-network designs, each run through a year of compounding disruption, and the exact stress level at which service breaks.

The agentic enterprise

Medium ↗

The Delegation Cliff

90,880 simulated designs of an AI-delegating organization. At a 95% quality floor they degrade gradually; at 99% the same design space develops a cliff.

Oversight & accountability

The machinery behind the studies

Several experiments run on AGENESIS and the other systems profiled in the workshop.

AGENESIS profile All experiments

Unpublished

From the notebook

Working pieces and explorations that live only on this site.

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top