Campaigns ·

Sprawl Cost Under a Fixed Monitoring Budget

Mixed 256 simulation units on held-out seeds

With the monitoring budget held flat, sprawl cost grows as fleet size to the power 1.4665 (95% interval 1.4101 to 1.5230); with per-agent monitoring it is 0.9001.

Registration and verificationPre-registered calibration with seven frozen hypotheses and a kill condition. Independent verification found 95 duplicate pairs (190 of 256 rows; 161 distinct simulations); one verdict moved from supported to refuted and is published as refuted.

The question

The governance literature says an agent portfolio you cannot retire from fast enough costs more than proportionally as it grows. Nobody had put a number on it. This campaign asked how fast the cost of a simulated agent fleet grows with its size, and whether the answer depends on how the monitoring budget is set.

What was registered

Seven hypotheses with frozen thresholds, written before the two seeds that count had ever run. The kill condition came from the review that commissioned the work: if cost turned out linear in agent count, the design was falsified and the project would retire.

  • H1: cost grows faster than linear under a monitoring budget that stays flat, with the interval's lower bound above 1.10.
  • H4: the exponent flattens at large sizes, tested by asking that the intervals on the four smallest and four largest sizes not touch.
  • H6: each of the four design factors moves the exponent by at least 0.10.

What was measured

Agents ramp up, drift silently out of spec, do damage while nobody notices, get flagged, queue for review and get retired. Cost is the harm from degraded agents still running plus staff time spent reviewing, and no cost term depends on how many agents exist, so any growth beyond linear has to come through the queue, the detection lag or the overlap between agents. 256 runs covered eight portfolio sizes from 57 to 787 agents, over an 18-month horizon.

What it found

Under a flat monitoring budget cost grows as fleet size to the power 1.4665, with a 95% interval of 1.4101 to 1.5230. Doubling the fleet multiplies cost by about 2.76. Where every agent carries its own check and review grows with the fleet, the exponent is 0.9001, about 1.87 per doubling. The mechanism shows in the detection lag: with a fixed inspection budget, the time from an agent going bad to someone noticing runs from 2.78 ticks (one tick is one week) at 57 agents to 22.49 ticks at 787.

Review capacity did nothing under diluted monitoring. In 55 of 64 matched cases, changing it produced a bit-identical simulation, because nothing had been flagged for the reviewers to look at. Detection sits upstream of review.

The independent verification found that 95 duplicate pairs (190 of the 256 rows, 161 distinct simulations), which the campaign's own guard had been scoped too narrowly to catch. Correcting for it left every point estimate in place and widened the intervals. That flipped H4 from supported to refuted, and the piece reports it as refuted. H6 was refuted too: two of the four factors moved the exponent by 0.062 and 0.089, under the frozen 0.10, so the follow-on is a two-factor study.

Headline numbers, as published
Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust)1.4665 [1.4101 to 1.5230]
Sprawl cost exponent in fleet size, per-agent monitoring with review that scales0.9001 [0.8664 to 0.9337]
Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring2.76× against 1.87×
Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical55 of 64
Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents2.78 to 22.49
Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145)1.4949 against 1.3238 [1.4037 to 1.5862; 1.2294 to 1.4182]
Banked runs the independent verification found to be duplicates95 duplicate pairs (190 of 256 rows; 161 distinct simulations)
1.21.31.41.51.6Smallest four, as reportedSmallest four, as reported: 1.4949 [1.4155, 1.5744]1.495Largest four, as reportedLargest four, as reported: 1.3238 [1.2618, 1.3858]1.324Smallest four, correctedSmallest four, corrected: 1.4949 [1.4037, 1.5862]1.495Largest four, correctedLargest four, corrected: 1.3238 [1.2294, 1.4182]1.324
Cost exponent fitted on the four smallest and the four largest portfolio sizes under diluted monitoring, with 95% intervals as first reported and after clustering on the result signature. The frozen test needed the intervals to stay apart. Rebuilt from the numbers in Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics.
Show the numbers as a table
FitExponent95% interval
Smallest four, as reported1.49491.4155 to 1.5744
Largest four, as reported1.32381.2618 to 1.3858
Smallest four, corrected1.49491.4037 to 1.5862
Largest four, corrected1.32381.2294 to 1.4182

What it does not claim

  • A calibration on one simulated industry, one company size, one degradation rate and one 18-month window. It was registered as a calibration that cannot confirm anything on its own.
  • The value 1.47 is an estimate awaiting a confirmatory study. The shape of the curve carries further than the number.
  • The flattening at large sizes is an observed direction with a failed test. The bar was left where it was.

Read the pieces

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top