Self-assessment · 30 questions · about ten minutes
The Delegation Index, live
Answer thirty questions about how your organisation hands work to AI agents and how it checks that work. The page maps your answers onto the five components of the lab's Delegation Index and puts your reading beside the lab's own Q3 2026 reading, 47.2, band 40.8 to 53.2.
This page runs what you report through a mapping rule it chose. The lab's figure comes from measured campaigns, simulations and controlled model runs on generated tasks, and nobody has calibrated the two against each other, so use the comparison to orient yourself. Your answers stay in this browser unless you export them.
Lab reading: Delegation Index v1, registered 2026-10-05, Q3 2026 reading published in The Measured Enterprise 2026, 7 Oct 2026Load a file this page exported and its answers and confidence marks fill the form. Anything that fails the checks is skipped and listed below. Loading replaces the answers now in the form.
Older draft found
Your reading
Five components on a ruler, beside the lab's Q3 2026 values
| Component | Answered | Your score | Your band | Lab Q3 2026 | Lab band |
|---|
Confidence, reported apart from the answers
| Component | Band from answers | Widened for confidence | Fairly sure | Guessing |
|---|
What moves the needle
This is arithmetic on your own answers under rule dil-1. For each component it finds the single answer whose change by one step shifts the composite furthest. Which practice deserves changing is a judgement the formula cannot make for you.
| Component | Answer | One step | Component moves | Composite moves | Ties |
|---|
Under this rule every straight-line item is worth 25 points, split across the answered items in its component and then across five components, so ties are common. S1 has unequal steps and can break them.
Export
The file carries your answers, confidence marks and reading along with the rule version, the question set and the lab's reference values. The only date in it is the day you export. A label helps when you compare two readings later; it goes into the file as typed, so leave names of people and organisations out of it.
Pooling readings across organisations
A pooled quarterly figure needs a collection step, somewhere that receives exports, checks and de-duplicates them, and prints the count behind any average. This page runs in your browser and has none.
If you want your reading counted once pooling starts, send the exported JSON to [contact placeholder: a public form or address to be added]. Pooling waits until there are enough readings to publish with a count. The lab's own index will keep reading its measurements alone.
Compare two readings
Two exports on the same rulers
Load two files this page exported: two teams, or one team a quarter apart. The page recomputes each reading from the answers inside the file, puts both on the rulers and lists B minus A. Nothing here touches the form above.
Reading A (filled circle, upper band)
Reading B (diamond, lower band)
| Component | A | B | B minus A | Bands |
|---|
Methods
How the page works
From answers to a component score
Each question has five answers and a "not sure" option. Answer 5 always describes the high end of that component's scale and answer 3 describes its balance point, so all items but one map on a straight line:
item score = (answer − 1) × 25 1 → 0, 2 → 25, 3 → 50, 4 → 75, 5 → 100
(all questions except S1)
S1 score = 50 × (2 − β), β = log2(m)
m = growth of oversight work when the fleet doubles: 4, 3, 2, 1.5, 1
1 → 0, 2 → 20.75, 3 → 50, 4 → 70.75, 5 → 100
component = mean of the answered item scores in that component
(a component needs at least 3 of its 6 questions answered)
This mapping is a design choice. Equal steps of 25 assume the five answers sit evenly along the component's scale, which nobody has measured. Every question weighs the same inside its component, again by choice. "Not sure" drops the question from the mean.
S1 is the one question with a direct formula. Its answers name how much oversight work grows when the number of agents doubles, which fixes a cost exponent, and the lab's sprawl formula turns that exponent into a score. No other answer set pins down the lab's quantity that exactly, so the rest keep the even 25-point steps, the simplest rule that puts answer 3 on the balance point.
Quantity items and practice items
A quantity item asks about the component's own quantity, for instance what share of skipped steps an audit finds. A practice item asks about a habit that drives that quantity, such as keeping a step list or a register of agents. Practice items stand in for a quantity the organisation rarely measures. They carry the same weight as quantity items in this rule.
| Item | Kind | Question |
|---|
Your bands and your composite
Your band for a component is a trimmed range of its item scores. With 5 or 6 answers the page sets aside the single lowest and single highest item and takes the lowest and highest of what is left. With 3 or 4 answers it takes the full range. The band is then stretched, if needed, to include the component score itself.
The composite follows the lab's own rule. Its point is the mean of the five component scores with equal weights. Its band is the mean of the five lower ends and the mean of the five upper ends. The composite needs all five components read.
A band here shows how far your answers inside a component disagree with each other. Read it as a range of disagreement; a confidence interval would need a sampling model this page lacks. When the composite band contains 50, your answers cannot tell balance from imbalance.
How confidence widens a band
Each question carries an optional mark: sure, fairly sure or guessing. The mark never changes your score. It widens the band, and the page reports the widened band in its own table and as a dashed outline on the ruler.
step allowance sure 0, fairly sure ½ step, guessing 1 step (unmarked counts as sure)
item range score at (answer − allowance) to score at (answer + allowance),
held inside answers 1 to 5; S1 interpolates between its five anchors
low spread item score − bottom of its range
high spread top of its range − item score
widened band band low − mean low spread to band high + mean high spread,
means over the component's answered items, held inside 0 to 100
composite mean of the five widened lower ends and of the five widened upper ends
Half a step on a straight-line item is 12.5 points and a full step is 25. A guess at answer 5 widens only downward, since the scale stops there. The allowances are a design choice, sized so that a guess spans the neighbouring answers on either side.
What moves the needle, and the answer-pattern notes
For each component the page tries every answered item one step up and one step down, recomputes the component mean and divides the change by five for the composite. It reports the largest move by size. Ties go to a step up, then to the earlier question. "Not sure" items stay out, since answering one would change how many items share the mean.
Three answer patterns trigger a note once ten or more questions carry a numbered answer: every answer a 3, every answer the same number, and answers that repeat a fixed cycle of two to five values in question order. A fourth note appears when fifteen or more questions are marked not sure. Each note says what the pattern does to the reading and leaves the scores as they are.
The lab's reading and where its numbers come from
The lab's Delegation Index reads one published number per component, the best measured design with oversight spend held fixed, puts each through a formula whose 50 is that component's balance point, and averages the five with equal weights. This page recomputes it from the published inputs when it loads.
| Component | Published input | Formula | Score | Band | Source |
|---|
Delegation price has no published band and enters at its point value at both ends, so the lab's band is narrower than the truth by that component's width. The band is the mean of band ends and carries no coverage level.
Appendix: one component recomputed, step by step
Pick a component and the page prints every step from the published inputs to the score that sits on the lab's ruler. The arithmetic runs live in your browser from the same inputs as the table above, and the self-check confirms the last line matches.
Why the two readings are different instruments
- The lab reads five numbers from its own synthetic campaigns: simulations and controlled model runs on generated tasks. Firms and markets sit outside what it reads.
- This page reads what one person reports about one organisation. Self-report drifts toward how people hope things work.
- The questions ask about the same five quantities in everyday terms. The page has not checked that an answer of 3 on a question corresponds to a measured 50 on the lab's scale. That calibration would need organisations whose oversight had been measured as well as surveyed.
- 50 marks where the two sides of a component cancel on its own definition. Whether to delegate stays your decision.
Drafts, imports, privacy and the self-check
Answers and confidence marks save as a draft in this browser's local storage after each change, under the key di-live-draft-v2. The draft records the question set it belongs to, . If a draft from another question set turns up, or one saved before the page stamped question sets, the page says so and asks before restoring it. If storage is blocked or full, the page keeps working and says the draft will not survive a reload. "Clear all answers" removes the draft.
Imports accept files up to 2 MB. An answer must be a whole number from 1 to 5, "not sure" or empty; anything else is skipped and listed. Keys that match no question, repeated keys and a stored reading that disagrees with its own answers each produce a note. Text from a file is shown as plain text.
Open the browser console to see the self-check. It runs fixed answer vectors through every computation on the page and compares the results with values worked out by hand, and it recomputes the lab's Q3 2026 reading from the published inputs.