Campaigns ·

Delegation Cliff

Mixed 90,880 configurations, zero error units

Each added level of delegation depth costs 1.50 on the quality margin; reviewer capacity buys back the most at +1.46.

Registration and verificationPredictions registered before the sweep; the two that failed were that observability outranks reviewer capacity and that an audit death spiral forms. A 200-pair revalidation certificate passed.

The question

Every organisation deploying AI agents is betting that pushing work further from human review is cheaper than keeping it close. This campaign priced that bet: where does delegation stop paying, and which design choices buy the margin back?

What was registered

Predictions registered before the sweep.

  • The design space has a cliff, a sharp boundary between robust and fragile configurations.
  • Better observability outranks raw reviewer capacity.
  • An audit death spiral, where eroding trust raises audit load until capacity saturates and quality falls.

What was measured

A simulated agent organisation with twelve design dimensions, including delegation depth, reviewer capacity, task coupling, model capability and self-check calibration. 90,880 configurations were each scored on the margin between the quality the organisation delivers and the floor it must hold. The campaign finished with zero error units. A revalidation certificate reran 200 random configurations from scratch: median difference 0.002, 95th percentile 0.0117, against a tolerance of 0.05.

What it found

Delegation depth carries the largest weight in the space, and it is negative: each increment costs 1.50 on the margin. Reviewer capacity buys back the most at +1.46. Model capability (+0.95) and self-check calibration (+0.92) follow, close enough to read as tied. Task coupling is the second-largest cost at −0.96.

The quality floor decides the failure geometry. At a 0.95 floor, 60% of the space is a broad transition band with no cliff. At 0.99, 79% of the space is fragile and the cliff appears. The observability prediction failed: observability complements reviewer capacity and does not replace it. The spiral prediction failed: mean margin rises across trust-response quintiles, including in the high-stress stratum.

Headline numbers, as published
Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948)−1.50
Change in quality margin per increment of reviewer capacity, the largest positive lever+1.46
Model capability against agent self-check calibration (read as tied)+0.95 against +0.92
Share of the design space in a broad transition band at a 0.95 quality floor (no cliff)60%
Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep)79%
Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.050.0117
−1.5−1.0−0.50.00.51.01.5Delegation depthDelegation depth: −1.499−1.499Reviewer capacityReviewer capacity: 1.4561.456Task couplingTask coupling: −0.958−0.958Model capabilityModel capability: 0.9490.949Self-check calibrationSelf-check calibration: 0.9150.915Rework costRework cost: −0.803−0.803Verification depthVerification depth: 0.4130.413Verification coverageVerification coverage: 0.3080.308Detection lagDetection lag: −0.259−0.259Workload volatilityWorkload volatility: −0.147−0.147Trust responseTrust response: 0.1020.102Escalation latencyEscalation latency: −0.007−0.007
What moves the delegation margin: linear-probe coefficients on the Sobol sample, band-only, n = 4,948. Positive buys margin, negative spends it. Rebuilt from the numbers in Pricing Agent Autonomy.
Show the numbers as a table
Design dimensionCoefficient
Delegation depth−1.499
Reviewer capacity1.456
Task coupling−0.958
Model capability0.949
Self-check calibration0.915
Rework cost−0.803
Verification depth0.413
Verification coverage0.308
Detection lag−0.259
Workload volatility−0.147
Trust response0.102
Escalation latency−0.007

What it does not claim

  • The coefficients are descriptive slopes on the Sobol sample inside the transition band, n = 4,948. They describe how the margin moves where a boundary exists. They are not causal estimates over the whole space.
  • The 0.99 result rests on one 256-unit calibration sweep. The 0.95 result is replicated on the 8,192-unit Sobol set. The two are not equally evidenced.
  • Reviewer capacity is a parameter in this design. Reviewers do not tire or queue.

Read the pieces

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top