# Dinand Tinholt — full site text Source: https://dinand.com/ · Generated 2026-10-08 · Guide: https://dinand.com/llms.txt ## Essays Longer thinking about AI and the people who have to live with it. History gets pulled in whenever it helps. ### The symbiotic enterprise URL: https://dinand.com/essays/the-symbiotic-enterprise/ Type: Working draft · 4 min read How human judgment and machine intelligence can grow more capable together. An enterprise changes when people begin to think with AI. A planner can explore a supply disruption while it is still developing. A finance team can ask how a shift in payment behavior will affect cash. Each exchange gives someone a chance to examine an assumption that might otherwise stay buried in a spreadsheet. #### The quality of the relationship The value of these exchanges depends on how work is organized. People need to understand where a recommendation came from and which parts of the decision remain uncertain. The system needs enough context to recognize when a technically plausible answer will fail in practice. Consider a supply planner deciding whether to move stock between distribution centers. An AI system can compare transport costs, service levels and inventory exposure. The planner may know that a customer has just changed its priorities. A useful workflow allows that knowledge to change the analysis before anyone commits to the shipment. #### Learning through use Every override offers a question worth investigating. Did the data arrive late? Was a constraint missing? Did a person bring knowledge the system could not access? Recording the reason makes the next decision easier to assess. Over time, these exchanges can improve both the model of the business and the way people use it. This is the starting point for a symbiotic enterprise: work designed so that human judgment and machine capabilities develop through repeated, visible interaction. It requires clear responsibility for decisions, time to learn from outcomes and enough curiosity to revisit the process itself. ### The past as a compass URL: https://dinand.com/essays/the-past-as-a-compass/ Type: Working draft · 4 min read Old maps and new technologies invite us to examine the assumptions behind what we see. An old map records a way of understanding the world. Its blank spaces, crowded coastlines and choices of scale reveal what its makers could see and what mattered to them. Reading it today involves noticing those choices as much as following its roads. #### The choices inside a model A business model also begins with choices. Someone decides which events to measure and which relationships to preserve. A forecast turns a complicated environment into a manageable picture. AI extends our ability to draw such pictures, and makes it easier to overlook the decisions that shaped them. History encourages a useful habit: ask how the picture was made. A beautiful map can place a city incorrectly. A fluent explanation can make an unsupported assumption feel settled. In either case, confidence grows more useful when we can inspect what supports it. #### Room for curiosity Books, art and philosophy offer other ways to notice what a familiar frame leaves out. A different vocabulary can change the question a team asks about technology. It can also make room for consequences that sit beyond the immediate business case. When a new tool promises a new world, I want to understand the older ideas traveling with it. That question gives the conversation somewhere useful to begin. ## Publications Everything I have published since 2022, sorted by year and theme. ### Self-healing master data URL: https://dinand.com/publications/self-healing-master-data/ Type: Working draft · 4 min read An executive brief on finding, proposing and verifying supply-chain data repairs. A supply chain can spend hours responding to a problem caused by one incorrect field. A wrong unit of measure changes the order quantity. An outdated lead time distorts a replenishment plan. A self-healing workflow aims to detect such errors, propose a repair and verify its effect within an agreed scope of authority. #### A practical example Suppose a supplier sells an item in cases of twelve, while a purchasing record treats each case as one unit. The workflow compares the supplier specification with purchasing history and receipts. It identifies the mismatch, shows the supporting records and proposes a conversion change. An authorized steward reviews changes that could affect open orders. After a correction, the workflow checks dependent calculations and records the previous value. If the result violates an agreed constraint, it restores the earlier state or escalates the case. That last step matters because a plausible local correction can still damage another process. #### Bound the authority Start with a narrow data domain and a well-defined failure pattern. Set rules for which repairs can proceed automatically and which require review. Track the rate of accepted repairs and recurrence of the underlying error. Include the cost of unnecessary corrections in the assessment. The operating goal is a repeatable loop with a clear owner. Each repair should leave enough evidence for someone to understand what changed, why it changed and whether the change helped. ### From dashboards to decisions URL: https://dinand.com/publications/from-dashboards-to-decisions/ Type: Working draft · 4 min read A working paper on bringing business context into an agentic decision workflow. A dashboard makes a change visible. The next task is to decide what to do about it. A decision intelligence workflow connects the signal to its likely causes, available responses and the person with authority to act. #### Following a question through the business Imagine that a finance leader asks why cash collection is below plan. The workflow checks the receivables data and identifies the customers driving the change. It can then examine disputes, payment patterns and operational events that may explain the delay. Each explanation needs a traceable source. Possible responses should be evaluated against the same business constraints. Accelerating a collection might damage a customer relationship or divert a team from a more valuable task. The system should make those assumptions visible enough for the decision owner to challenge them. #### Close the loop Once an action is approved, the workflow should retain the decision and its expected effect. Comparing the outcome with that expectation creates evidence for the next recommendation. It also reveals whether an apparently successful action worked for the reason the system predicted. This working paper outlines a design direction. Its examples are illustrative and do not describe a measured client implementation. ### Skin in the Game Made the AI Supervisor Worse URL: https://dinand.com/publications/skin-in-the-game/ Original: https://www.linkedin.com/pulse/skin-game-made-ai-supervisor-worse-dinand-tinholt-13jic?trk=public_post Published: LinkedIn, 2026-09-09 · Topic: AI oversight A controlled experiment on the cost of wrongful blocks and the incentives that shape an AI reviewer. Across 30,000 synthetic decision episodes, a reviewer carrying a record of past harm became more likely to block correct decisions. The article examines the consequences for oversight design and argues for measuring the opportunity cost of refusal. It also describes the study’s limits: one model family, constructed tasks and results that should be interpreted within that setting. ### Inventory Optimization Under Uncertainty: Broad random shocks trained a tougher inventory planner than a clever AI adversary did URL: https://dinand.com/publications/inventory-under-uncertainty/ Original: https://medium.com/@tinholt/inventory-optimization-under-uncertainty-broad-random-shocks-trained-a-tougher-inventory-planner-396134e2c5e3 Published: Medium, 2026-06-15 · Topic: Supply-chain resilience An inventory-planning experiment explores which kinds of stress prepare a policy for unfamiliar disruptions. A simulated distribution center becomes a laboratory for testing inventory policies against supply disruptions. A language model manages the experimental loop. Broad random shocks produced a directional advantage over targeted adversarial training, with the article explicitly noting that five runs do not establish statistical significance. The documented research process and its self-audit are central to the work. ### Move 37: How Agentic AI Turns Back-Office Drudgery into Bold Process Breakthroughs URL: https://dinand.com/publications/move-37/ Original: https://www.capgemini.com/insights/expert-perspectives/move-37-how-agentic-ai-turns-back-office-drudgery-into-bold-process-breakthroughs/ Published: Capgemini, 2026-04-29 · Topic: Process reinvention Why the larger opportunity in agentic AI is not polishing an inherited workflow, but questioning whether the workflow should exist in its current form. Using AlphaGo’s unexpected Move 37 as a metaphor, this Capgemini perspective argues for redesigning back-office work around new capabilities rather than using AI only for incremental efficiency. It focuses on a change in framing: begin with the outcome and the decisions involved, then ask what a process would look like if it were designed now. ### Building a Fake Economy Run by AI Agents to Negotiate Trade Promotions: Lessons from a CPG Simulation URL: https://dinand.com/publications/trade-promotion-agents/ Original: https://www.linkedin.com/pulse/building-fake-economy-run-ai-agents-negotiate-trade-lessons-tinholt-drldc?trk=public_post Published: LinkedIn, 2026-02-28 · Topic: Agents in supply chains Manufacturer and retailer agents negotiate promotions inside a simulated market with competing incentives. This article follows a trade-promotion experiment built around a 52-week market simulation. Two agents negotiate discounts and display fees, observe outcomes, and adjust their behavior. It explores how seasonal demand, inventory and reward design affect their choices. The work ran on two NVIDIA DGX Spark machines and presents simulation findings rather than results from a production deployment. ### How to Make AI Agents Accurate: Stop Treating Memory Like Chat History URL: https://dinand.com/publications/agent-memory-and-state/ Original: https://medium.com/@tinholt/how-to-make-ai-agents-accurate-stop-treating-memory-like-chat-history-40eb8e0ea437 Published: Medium, 2025-12-17 · Topic: Agent engineering A practical approach to keeping an agent’s constraints and current decisions clear during long sessions. This piece proposes managing explicit project state as a way to reduce drift in long-running agent sessions. It describes compact records of architectural decisions, constraints and current work, with periodic snapshots and resets. It also examines the risks of stale records and overly rigid rules. The article presents a working practice for building with agents. ### The Next Frontier of Agentic AI: Context That Learns URL: https://dinand.com/publications/context-that-learns/ Original: https://medium.com/@tinholt/the-next-frontier-of-agentic-ai-context-that-learns-3246d4e73744 Published: Medium, 2025-10-12 · Topic: Adaptive systems How feedback from an agent’s work can inform the context it uses on the next task. A reflection on agentic context engineering and the possibility of improving an agent through its working context. The article explores how experience could inform subsequent actions and considers the implications for enterprise processes. It asks what governance is needed when the instructions and knowledge surrounding a deployed model continue to change. ### GPT-5’s Multimodal Breakthrough: How AI Just Crossed the Expert Threshold in Business Reasoning URL: https://dinand.com/publications/multimodal-business-reasoning/ Original: https://www.linkedin.com/pulse/gpt-5s-multimodal-breakthrough-how-ai-just-crossed-expert-tinholt-ewyhc?trk=public_post Published: LinkedIn, 2025-09-03 · Topic: Multimodal AI A business perspective on research combining visual information with complex written evidence. The article considers what advances in multimodal medical reasoning might mean for enterprise work involving documents, images and structured information. It uses a research result as a starting point for discussing business applications. Those applications are the author’s interpretation; the underlying medical evaluation is not itself a validation of performance on enterprise tasks. ### The AI Capability Overhang: Why the Most Powerful Technology Isn’t Working (yet) URL: https://dinand.com/publications/ai-capability-overhang/ Original: https://medium.com/@tinholt/the-ai-capability-overhang-why-the-most-powerful-technology-isnt-working-yet-1855eec909be Published: Medium, 2025-06-02 · Topic: Enterprise adoption Why organizations need to develop the operating capacity to use the AI capabilities already available. The article examines the distance between model capability and business implementation. It discusses fragmented data, integration work and the practical demands of governance. Its proposed approach begins with a specific business process, supported by infrastructure and people who understand both the technology and the work it is meant to improve. ### Symbiotic Intelligence Theory (SIT): Rethinking Human Potential in the Age of AI URL: https://dinand.com/publications/symbiotic-intelligence-theory/ Original: https://medium.com/@tinholt/symbiotic-intelligence-theory-sit-rethinking-human-potential-in-the-age-of-ai-4664c8840652 Published: Medium, 2025-01-23 · Topic: Human–AI collaboration A framework for thinking about human agency and the ways AI can extend our capacity to think together. This foundational essay introduces Symbiotic Intelligence Theory as a proposed framework for human–AI collaboration. It considers human capabilities, AI-assisted thinking and more integrated forms of cooperation. Alongside the possibilities, the article explores cognitive dependence, unequal access and the need to preserve judgment as AI becomes part of daily intellectual work. ### From Chaos to Managed Complexity: Adopting an Ecosystem Mindset in Data & Analytics URL: https://dinand.com/publications/from-chaos-to-managed-complexity/ Original: https://www.linkedin.com/pulse/from-chaos-managed-complexity-adopting-ecosystem-mindset-tinholt-a3g6c Published: LinkedIn, 2024-11-01 · Topic: Data ecosystems A case for treating modern data environments as evolving ecosystems rather than trying to force every source and decision into one rigid center. This article considers how decentralized data, layered technology and growing organizational complexity change the work of data leadership. It argues for an ecosystem mindset that coordinates standards and outcomes while leaving room for local context, rather than equating control with complete centralization. ### The Liminality of Data: A Glimpse Between the Thresholds URL: https://dinand.com/publications/liminality-of-data/ Original: https://www.linkedin.com/pulse/liminality-data-glimpse-between-thresholds-dinand-tinholt Published: LinkedIn, 2023-10-03 · Topic: Change and uncertainty What the idea of liminality can teach data and AI leaders about working between an established order and one that has not yet taken shape. Prompted by the language of liminality, this reflection looks at periods when familiar structures are weakening but their replacements remain uncertain. It applies that lens to data and AI, asking how leaders can stay useful and curious while practices, roles and expectations are still being renegotiated. ### Creating a Data-Powered Culture URL: https://dinand.com/publications/creating-a-data-powered-culture/ Original: https://www.linkedin.com/pulse/creating-data-powered-culture-dinand-tinholt Published: LinkedIn, 2022-01-21 · Topic: Data culture A durable starting point for making data part of everyday decisions: connect technology, shared habits and organizational confidence. First published in Capgemini’s Data-Powered Innovation Review, this article treats data culture as an organizational practice rather than a technology installation. It considers the conditions that help people use evidence with confidence and make data part of ordinary work. ## Reading Notes What I took away from the books and papers that changed my mind about something. ### Ways of Seeing URL: https://dinand.com/reading-notes/ways-of-seeing/ Type: Reading-note draft · 2 min read John Berger, and the questions we bring to an image. An image arrives with a frame. Its placement, caption and surroundings help shape the way we read it. Berger’s title offers a useful starting question for anyone working with AI: what influences our interpretation before we begin to judge the content? #### A question for technology An AI interface also frames an answer. A confident paragraph, an authoritative voice or an elegant chart can affect how much scrutiny an idea receives. Looking carefully means attending to the presentation as well as the claim. A reading prompt: choose an AI-generated answer and consider how its meaning would change if it appeared as a rough notebook entry rather than a polished report. Which parts of your trust come from evidence, and which come from presentation? #### For the margin What would a reader need to see alongside an answer to evaluate it well? This is an original draft reflection prompted by the book’s themes, with no quotations or page-specific annotations. ### The Society of the Spectacle URL: https://dinand.com/reading-notes/society-of-the-spectacle/ Type: Reading-note draft · 2 min read Guy Debord, representation, and a question about synthetic media. Debord’s work invites a reading of social life through the representations that mediate it. For someone thinking about AI, that raises a contemporary question: how does the growing supply of synthetic content affect our relationship to the events it claims to show? #### The distance between image and event A generated image can travel with the visual confidence of a photograph. A synthetic account can feel like someone’s direct experience. As readers, we need ways to understand the connection between a representation and the world beyond it. A reading prompt: follow a claim from a social post to its source. Notice how the framing changes along the way. What was made more vivid? Which uncertainty disappeared? This is a draft reading lens for exploring the book alongside AI. It offers no direct quotations and does not claim that Debord anticipated the technical details of contemporary models. ## Research Simulations I run when I want to know if an idea about AI survives contact with real numbers. ### When monitoring stops scaling URL: https://dinand.com/research/when-monitoring-stops-scaling/ Type: Study summary · draft · 4 min read What a simulation suggests about the cost of overseeing a growing agent portfolio. In the AGENESIS-2 simulation, modeled portfolio cost grew by about 2.76 times each time the agent portfolio doubled while the monitoring team stayed the same size. That corresponds to a scaling exponent of 1.466. A design that gave each agent its own check grew roughly in proportion to the portfolio. #### What the result puts into focus The useful question concerns the capacity of the oversight design. If agents expand faster than the ability to check their work, adding agents changes the operating conditions for the whole portfolio. A budget based on cost per agent can miss that interaction. A CIO can use this as a reason to test several portfolio sizes before committing to a rollout. Measure the work required to review decisions alongside the consequences of missed errors. Examine whether the review process still works under the larger workload. #### Correction and limits The supplied executive brief describes 256 pre-registered simulation units. Independent verification identified 95 duplicate runs that differed only in a setting with no effect at most portfolio sizes. Correcting the arithmetic left the headline exponent at 1.466; one reported verdict changed. These are simulation findings from the supplied study summary. This page does not establish an empirical scaling law for deployed enterprise agents. The raw ledger and full methodology are not included here. The accompanying explorer simply illustrates the reported exponent; it does not forecast actual costs. ### When liability changes the reviewer URL: https://dinand.com/research/oversight-and-liability/ Type: Published article overview · 4 min read A controlled simulation of how a record of past harm affected an AI supervisor. The published account reports that liability pressure increased wrongful blocking in a controlled simulation. The share of worker errors a blocker could address fell from 93% at the reversible tier to 2% at the most irreversible tier. The original article includes the experiment’s design, findings and limits. Read it for the full account. Read the published article on LinkedIn ↗ ## Experiments Apps and research systems I built. Two of them run right here in your browser. ### Monitoring cost explorer URL: https://dinand.com/experiments/monitoring-cost-explorer/ Type: Interactive tool · Try it Explore how portfolio size changes modeled oversight cost. ### Decision brief builder URL: https://dinand.com/experiments/decision-brief-builder/ Type: Interactive tool · Try it Turn a question, evidence and constraints into a structured brief. ### AI Control Plane URL: https://dinand.com/experiments/ai-control-plane/ Status: Private working prototype · Separate service-backed application A register and evidence layer for examining an enterprise portfolio of AI agents: who owns them, how far they act, and what supports the trust placed in them. Question behind the system: As agents multiply, the hard problem shifts from building one capable system to understanding the estate as a whole. Capabilities: - Map agents, ownership, autonomy and operating context - Keep evidence and qualifications attached to every assessment - Preserve snapshots, signals and an audit history Public boundary: This page describes the design at a high level. The working build is private and service-backed; no internal materials, client information, credentials or operational controls are exposed here. ### AGENESIS URL: https://dinand.com/experiments/agenesis/ Status: Completed research campaign · Research case study An agent-based laboratory for exploring how different levels of AI participation interact with enterprise structure, market conditions and industry context. Question behind the system: An enterprise does not adopt AI in isolation; decisions propagate through competitors, suppliers, customers and the wider system. Capabilities: - Model connected enterprises across multiple industry settings - Sweep organization, market and agentification parameters - Turn a large simulation campaign into an opportunity atlas Public boundary: The campaign captured run configurations and duration, but final per-enterprise financial metrics were not persisted. This profile therefore describes the experimental system and does not claim financial outcome findings. ### SCENARIO-X Studio URL: https://dinand.com/experiments/scenario-x/ Status: Working local application · Separate FastAPI application An interactive workbench for designing multi-product supply-chain experiments and inspecting where average performance hides a weak tail. Question behind the system: A portfolio can look healthy in aggregate while one product, scenario or shared constraint is already failing. Capabilities: - Design product portfolios, disruptions and policy grids - Compare cost-service frontiers and weakest-SKU outcomes - Diagnose binding constraints and export experiment artifacts Public boundary: The interface depends on a local simulation service and persisted run workspace. The public site provides an introduction; a deployed instance can be linked separately when its service boundary is ready. ### Assay Bench URL: https://dinand.com/experiments/assay-bench/ Status: Working research product · Separate Next.js application A workbench for measuring the quality and checking cost of AI outputs when there is no simple answer key. Question behind the system: The cost that constrains an AI process may not be generation. It may be the cost of becoming confident enough to trust the output. Capabilities: - Report quality with uncertainty and explicit blind spots - Plan samples against both confidence and checking cost - Keep judged cases, methods and human effort inspectable Public boundary: The application sits beside a process rather than controlling it. This public profile includes no seed database, credentials, model endpoints, client material or live review workflow. ### Personal Quant Agent URL: https://dinand.com/experiments/personal-quant-agent/ Status: Paper-only research prototype · Read-only project profile A research system that scans public market data, tests a small strategy library, applies explicit risk controls and records simulated decisions in a paper portfolio. Question behind the system: A useful trading experiment should make weak evidence, regime dependence, costs and failure conditions more visible—not hide them behind an autonomous agent. Capabilities: - Backtest signals net of modeled fees and slippage - Challenge strategies with walk-forward and Monte Carlo checks - Apply paper-portfolio sizing, stops and circuit breakers Public boundary: This is a personal research prototype, not investment advice or a claim of expected performance. It has no live order routing, and the public site exposes no portfolio state or trading controls. ## Side Projects Things I build for the fun of it, like a downtown Chicago you can walk through at any hour. ## About Born in the Netherlands, at home in Chicago, with a lifelong habit of asking one more question. Dinand Tinholt is an AI consultant, researcher, writer and builder. Originally from the Netherlands; now based in Chicago. - LinkedIn: https://www.linkedin.com/in/tinholt - Medium: https://medium.com/@tinholt - GitHub: https://github.com/dtinholt - Threads: https://www.threads.net/@tinholt - Instagram: https://www.instagram.com/tinholt/ ## Complete article archive ### The Open Decision Head: An Executive Briefing Published: 2026-10-07 · LinkedIn: https://www.linkedin.com/feed/update/urn:li:activity:7513629422663434240/ Theme: The agentic enterprise A retrained open decision head beat both a rule table and the hosted service on 800 fresh cases, after a first version that did not. ### A Trial Balance for Agent Omissions Published: 2026-10-01 · Medium: https://medium.com/@tinholt/a-trial-balance-for-agent-omissions-568a41c59aa1 Theme: Oversight & accountability Agent logs record what happened and have no row for what didn’t. A pre-registered test of commitment accounting, an obligation ledger reconciled against executed calls, for catching omissions in procurement-workflow traces. ### Jev is a new model from TypeSafe AI that helps software make small, specific decisions. Published: 2026-09-27 · LinkedIn: https://www.linkedin.com/posts/tinholt_jev-is-a-new-model-from-typesafe-ai-that-share-7510004413076250624-2pCZ/ Theme: Models & the AI market Pre-announcement: the benchmark, code and pilot guide follow once the results are verified. ### Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics Published: 2026-09-17 · Medium: https://medium.com/@tinholt/sprawl-cost-is-superlinear-under-a-fixed-monitoring-budget-a-pre-registered-calibration-of-agent-e24cbaa7e62a Theme: Oversight & accountability AGENESIS-2 adds a portfolio layer to the AGENESIS enterprise simulation to measure whether agent sprawl costs more than proportionally under a fixed monitoring budget. ### Your AI Review Team May Be Waiting for Failures It Cannot See Published: 2026-09-17 · LinkedIn: https://www.linkedin.com/feed/update/urn:li:activity:7506213115978489856/ Theme: Oversight & accountability The executive reading of the sprawl-cost result. ### Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers Published: 2026-09-10 · Medium: https://medium.com/@tinholt/accountability-makes-oversight-worse-a-pre-registered-test-of-liability-exposed-ai-supervision-ef4fdf46bad3 Theme: Oversight & accountability A pre-registered test of whether making an AI supervisor answerable for past harm improves its oversight, and why it backfired across irreversibility tiers. ### Skin in the Game Made the AI Supervisor Worse Published: 2026-09-09 · LinkedIn: https://www.linkedin.com/pulse/skin-game-made-ai-supervisor-worse-dinand-tinholt-13jic Theme: Oversight & accountability A controlled experiment on the cost of wrongful blocks and the incentives that shape an AI reviewer. ### Pricing Agent Autonomy Published: 2026-09-03 · Medium: https://medium.com/@tinholt/pricing-agent-autonomy-0a2bb5b97caf Theme: Oversight & accountability A simulation sweeping twelve dimensions of organisational design to price the bet that moving work further from human review is cheaper. Delegation depth carries the strongest weight, and it is negative. ### When the Second Opinion Shares the Blind Spot Published: 2026-09-02 · Medium: https://medium.com/@tinholt/when-the-second-opinion-shares-the-blind-spot-7e4443dcb081 Theme: Oversight & accountability One model on two machines, with different quantisation and serving stacks, made the same mistakes. What separates one reviewer from another is the model family: a case for dissimilar redundancy. ### A ledger for agentic decisions Published: 2026-08-28 · Medium: https://medium.com/@tinholt/a-ledger-for-the-decisions-nobody-priced-625c452e43fd Theme: Oversight & accountability Difficulty is a property of the work; risk is a property of the consequences. Testing, in simulation and as a product, the rule that decides which agentic decisions a person must see. ### The Breaking Point Published: 2026-08-27 · Medium: https://medium.com/@tinholt/the-breaking-point-2b68dbd7d7c2 Theme: The agentic enterprise 45,312 simulated supply-network designs, each run through a year of compounding disruption, and the exact stress level at which service breaks. ### The Delegation Cliff Published: 2026-08-16 · Medium: https://medium.com/@tinholt/the-delegation-cliff-ed787b1ebe22 Theme: Oversight & accountability 90,880 simulated designs of an AI-delegating organization. At a 95% quality floor they degrade gradually; at 99% the same design space develops a cliff. ### AI Agents May Have a Reason to Trust Each Other Published: 2026-08-09 · Medium: https://medium.com/@tinholt/ai-agents-may-have-a-reason-to-trust-each-other-3ba117adc87a Theme: The agentic enterprise On a paper led by Alexander Meulemans: why foundation-model agents may cooperate in a one-shot Prisoner’s Dilemma, and what that means for large populations of agents. ### Nobody Slows Down First Published: 2026-08-06 · Medium: https://medium.com/@tinholt/nobody-slows-down-first-2edd34ecaf97 Theme: Models & the AI market Six days apart, the same industry asked Washington to keep AI capability spreading and to build the machinery for holding it back. ### Inventory Optimization Under Uncertainty: Broad random shocks trained a tougher inventory planner than a clever AI adversary did Published: 2026-06-15 · Medium: https://medium.com/@tinholt/inventory-optimization-under-uncertainty-broad-random-shocks-trained-a-tougher-inventory-planner-396134e2c5e3 Theme: The agentic enterprise An inventory-planning experiment explores which kinds of stress prepare a policy for unfamiliar disruptions. ### Move 37: How Agentic AI Turns Back-Office Drudgery into Bold Process Breakthroughs Published: 2026-04-29 · Capgemini: https://www.capgemini.com/insights/expert-perspectives/move-37-how-agentic-ai-turns-back-office-drudgery-into-bold-process-breakthroughs/ Theme: The agentic enterprise Why the larger opportunity in agentic AI is not polishing an inherited workflow, but questioning whether the workflow should exist in its current form. ### Building a Fake Economy Run by AI Agents to Negotiate Trade Promotions: Lessons from a CPG Published: 2026-02-28 · Medium: https://medium.com/@tinholt/building-a-fake-economy-run-by-ai-agents-to-negotiate-trade-promotions-lessons-from-a-cpg-785384cdfb91 · LinkedIn: https://www.linkedin.com/pulse/building-fake-economy-run-ai-agents-negotiate-trade-lessons-tinholt-drldc Theme: The agentic enterprise Manufacturer and retailer agents negotiate promotions inside a simulated market with competing incentives. ### The Weights Got Free. The Grip Moved. Published: 2026 · Medium: https://medium.com/@tinholt/the-weights-got-free-the-grip-moved-fadf3d80fb4e Theme: Models & the AI market Kimi K3 put 2.8 trillion parameters on a public server. Reading that as decentralization gets the economics backwards. ### The Truth Ledger Published: 2026 · Medium: https://medium.com/@tinholt/the-truth-ledger-3e01dac4e8ca Theme: Ideas & culture Old Frames, New Machines · An Essay in Nine… ### The Most Powerful AI Ever Released Is Probably Being Wasted Right Now Published: 2026 · Medium: https://medium.com/@tinholt/the-most-powerful-ai-ever-released-is-probably-being-wasted-right-now-67eb38b2ad53 Theme: Work, economy & the future There is an old Greek storyteller named Aesop (you may remember the tortoise and the hare) whose genius had nothing to do with animals. ### The Vertical AI Playbook Is Here, And Legal Is Just The Latest Move Published: 2026 · Medium: https://medium.com/@tinholt/the-vertical-ai-playbook-is-here-and-legal-is-just-the-latest-move-65b71e4a7f5f Theme: Models & the AI market A week before Anthropic announced Claude for Legal, it announced Claude for Financial Services. ### Three bets on where the agent lives Published: 2026 · Medium: https://medium.com/@tinholt/three-bets-on-where-the-agent-lives-4c6f1c906383 Theme: Models & the AI market In a single quarter, three of the most important enterprise software vendors placed three competing claims on the agent… ### The AI Infrastructure Layer: Why Runtime Ownership Matters More Than Model Choice Published: 2026 · Medium: https://medium.com/@tinholt/the-ai-infrastructure-layer-why-runtime-ownership-matters-more-than-model-choice-1edf8e097603 Theme: Models & the AI market Enterprise technology teams are debating which AI model to use. GPT or Claude? Gemini or a fine-tuned open source model? ### What Happens When One Company Adopts AI Before Its Competitors? Published: 2026 · Medium: https://medium.com/@tinholt/what-happens-when-one-company-adopts-ai-before-its-competitors-cd976e5a8dba Theme: The agentic enterprise Results from a 100-run multi-agent enterprise simulation across four… ### AI-Driven Inventory Optimization Under Deep Uncertainty Published: 2026 · Medium: https://medium.com/@tinholt/ai-driven-inventory-optimization-under-deep-uncertainty-607fd80386d1 Theme: The agentic enterprise An Autoresearch… ### The Compounding Gap Published: 2026 · Medium: https://medium.com/@tinholt/the-compounding-gap-295f2dc6cf08 Theme: Work, economy & the future ### Twenty-one releases in under three months. Published: 2026 · Medium: https://medium.com/@tinholt/twenty-one-releases-in-under-three-months-7de36a1a1cd0 Theme: Models & the AI market Since January 2026, Anthropic has shipped two new foundation models (Opus 4.6 and Sonnet 4.6), and nineteen distinct product launches… ### Your AI Agent’s Brain: Specialized or Well-Briefed? Published: 2026 · Medium: https://medium.com/@tinholt/your-ai-agents-brain-specialized-or-well-briefed-88ec339c6db4 Theme: The agentic enterprise What a Controlled Experiment Reveals About Enterprise AI… ### Two Reports, One Question: What Does AI Actually Mean for Your Future? Published: 2026 · Medium: https://medium.com/@tinholt/two-reports-one-question-what-does-ai-actually-mean-for-your-future-9689d6bceec6 Theme: Work, economy & the future In the span of a few days in late February 2026, two documents landed in the financial and technology discourse with sharply… ### 10x the Impact of the Industrial Revolution, and We’re Just Getting Started Published: 2026 · Medium: https://medium.com/@tinholt/10x-the-impact-of-the-industrial-revolution-and-were-just-getting-started-08e90b8e3659 Theme: Work, economy & the future A year ago, a task that took a full year now takes three days. ### Why Industry Is the Only Thing That Separates AI Theater from AI Transformation Published: 2026 · Medium: https://medium.com/@tinholt/why-industry-is-the-only-thing-that-separates-ai-theater-from-ai-transformation-b1602bff12c5 Theme: The agentic enterprise I’ll save you the conference ticket and the keynote. ### Agent Swarms Are Here Published: 2026 · Medium: https://medium.com/@tinholt/agent-swarms-are-here-4cea3c59d7fb Theme: Models & the AI market Three major model drops in ten days. Moonshot AI shipped Kimi K2.5 on January 27th with a built-in agent swarm. ### How to Make AI Agents Accurate: Stop Treating Memory Like Chat History Published: 2025-12-17 · Medium: https://medium.com/@tinholt/how-to-make-ai-agents-accurate-stop-treating-memory-like-chat-history-40eb8e0ea437 Theme: Data, context & memory A practical approach to keeping an agent’s constraints and current decisions clear during long sessions. ### The Next Frontier of Agentic AI: Context That Learns Published: 2025-10-12 · Medium: https://medium.com/@tinholt/the-next-frontier-of-agentic-ai-context-that-learns-3246d4e73744 Theme: Data, context & memory How feedback from an agent’s work can inform the context it uses on the next task. ### GPT-5’s Multimodal Breakthrough: How AI Just Crossed the Expert Threshold in Business Reasoning Published: 2025-09-03 · Medium: https://medium.com/@tinholt/gpt-5s-multimodal-breakthrough-how-ai-just-crossed-the-expert-threshold-in-business-reasoning-9a0e24ce9b29 · LinkedIn: https://www.linkedin.com/pulse/gpt-5s-multimodal-breakthrough-how-ai-just-crossed-expert-tinholt-ewyhc Theme: Models & the AI market A business perspective on research combining visual information with complex written evidence. ### Turning Data Exhaust into Business Intelligence: The Hidden Value in Your Enterprise Operations Published: 2025-08-14 · Medium: https://medium.com/@tinholt/turning-data-exhaust-into-business-intelligence-the-hidden-value-in-your-enterprise-operations-e55356df99ca · LinkedIn: https://www.linkedin.com/pulse/turning-data-exhaust-business-intelligence-hidden-value-tinholt-qvxnc Theme: Data, context & memory In the landscape of enterprise AI, we often get caught up in the sheer volume of data. But the real secret isn’t just having mountains of information. It’s learning to recognize the patterns that live within it. ### From Language Models to World Models: Building the AI Operating Layer for Business Published: 2025-08-10 · Medium: https://medium.com/@tinholt/from-language-models-to-world-models-building-the-ai-operating-layer-for-business-c09e9be33236 · LinkedIn: https://www.linkedin.com/pulse/from-language-models-world-building-ai-operating-layer-dinand-tinholt-ztikc Theme: The agentic enterprise Everyone is staring at the same leaderboard again. GPT-5 here, a new benchmark there, another round of debates about who is ahead in model scale. The conversation is familiar and comfortable. ### Are We Cracking the Code of AGI? A Tiny Brain-Inspired AI Just Made a Big Statement Published: 2025-08-04 · Medium: https://medium.com/@tinholt/are-we-cracking-the-code-of-agi-a-tiny-brain-inspired-ai-just-made-a-big-statement-014104535b82 · LinkedIn: https://www.linkedin.com/pulse/we-cracking-code-agi-tiny-brain-inspired-ai-just-made-dinand-tinholt-uzjyc Theme: Models & the AI market There’s something happening in AI research that deserves the full attention of business leaders. ### The Great Untangling: Why Business Leaders Must Move From Agentic Mush to an Agentic Mesh Published: 2025-08-01 · Medium: https://medium.com/@tinholt/the-great-untangling-why-business-leaders-must-move-from-agentic-mush-to-an-agentic-mesh-692602f95a86 · LinkedIn: https://www.linkedin.com/pulse/great-untangling-why-business-leaders-must-move-from-agentic-tinholt-e8b9c Theme: The agentic enterprise Across boardrooms and executive strategy sessions, one phrase keeps surfacing: “We need to do something with Agentic AI.” So the sprint begins. Teams deploy chatbots. Product leads automate decisions. ### The AI Supermodel Has Arrived. And No, It’s Not the End of Beauty. It’s the Start of a Smarter Industry. Published: 2025-07-28 · Medium: https://medium.com/@tinholt/the-ai-supermodel-has-arrived-4ca6fae231da · LinkedIn: https://www.linkedin.com/pulse/ai-supermodel-has-arrived-its-end-beauty-start-smarter-dinand-tinholt-ygdcc Theme: Retail & consumer goods Vogue. Long the cathedral of couture and the arbiter of aspirational aesthetics. And now? A bold new guest walks the hallowed pages of its August issue: Seraphinne Vallora’s AI-generated model for a Guess campaign. She’s blonde. ### You Don’t Have an Agent Problem. You Have an Observability Problem. Published: 2025-07-25 · Medium: https://medium.com/@tinholt/you-dont-have-an-agent-problem-you-have-an-observability-problem-2d2b2d2e42cc · LinkedIn: https://www.linkedin.com/pulse/you-dont-have-agent-problem-observability-dinand-tinholt-iorcc Theme: The agentic enterprise Agentic AI is becoming the next big promise in enterprise automation. These systems go beyond answering questions. They take actions, chain tasks, call APIs, interact with memory, and adapt over time. They are not static copilots. ### The Myth of the Single Source of Truth Published: 2025-07-21 · Medium: https://medium.com/@tinholt/the-myth-of-the-single-source-of-truth-1ad4d8625215 · LinkedIn: https://www.linkedin.com/pulse/myth-single-source-truth-dinand-tinholt-m3ccc Theme: Data, context & memory Companies have spent years chasing a single source of truth. The idea is simple. One clean, standardized, central repository. Everyone uses the same data to make better decisions. ### From Centaur to Enterprise: How AI is Learning to Think Like Us (and What That Means for Business) Published: 2025-07-14 · Medium: https://medium.com/@tinholt/from-centaur-to-enterprise-how-ai-is-learning-to-think-like-us-and-what-that-means-for-business-86113c4378cf · LinkedIn: https://www.linkedin.com/pulse/from-centaur-enterprise-how-ai-learning-think-like-us-dinand-tinholt-igtic Theme: The agentic enterprise In the TV series Westworld, a system called Rehoboam shaped the future by predicting human behavior with near-perfect accuracy. Fiction, of course—until it isn’t. ### You Can’t Automate Dysfunction Published: 2025-06-28 · LinkedIn: https://www.linkedin.com/pulse/you-cant-automate-dysfunction-dinand-tinholt-ys6yc Theme: The agentic enterprise The agentic AI hype machine is running full speed. Everywhere I look, someone is wiring an LLM into a brittle business process and expecting magic. Startups are doing it. Enterprises are doing it. Everyone is nodding along. ### The Real Failure Isn’t Agentic AI — It’s Corporate Imagination Published: 2025-06-27 · LinkedIn: https://www.linkedin.com/pulse/real-failure-isnt-agentic-ai-its-corporate-dinand-tinholt-s7e5c Theme: The agentic enterprise From the moment Gartner’s June 25 report landed, proclaiming that over 40 percent of agentic AI initiatives will be abandoned by 2027, journalists seized on the alarm: “rising costs,” “unclear business value,” and worries about… ### Reimagining the Consumer Products Value Chain with AI Published: 2025-06-27 · LinkedIn: https://www.linkedin.com/pulse/reimagining-consumer-products-value-chain-ai-dinand-tinholt-xn5kc Theme: Retail & consumer goods ### Context Engineering Will Decide the Winners in AI Published: 2025-06-26 · LinkedIn: https://www.linkedin.com/pulse/context-engineering-decide-winners-ai-dinand-tinholt-dfame Theme: Data, context & memory The conversation around AI inside most companies is still stuck on prompt engineering — crafting clever questions to get clever answers. But the real work of building effective AI systems isn’t about witty phrasing. ### The AI Capability Overhang: Why the Most Powerful Technology Isn't Working (yet) Published: 2025-06-02 · LinkedIn: https://www.linkedin.com/pulse/ai-capability-overhang-why-most-powerful-technology-isnt-tinholt-0k5fc Theme: The agentic enterprise Why organizations need to develop the operating capacity to use the AI capabilities already available. ### The Real AI Race Is Happening in Your Business Published: 2025-05-27 · LinkedIn: https://www.linkedin.com/pulse/real-ai-race-happening-your-business-dinand-tinholt-5wv4c Theme: The agentic enterprise While superpowers debate the future of intelligence, smart companies are deploying it… ### Enterprise Models: The Rise of the Thinking Organization Published: 2025-05-22 · LinkedIn: https://www.linkedin.com/pulse/enterprise-models-rise-thinking-dinand-tinholt-7evdc Theme: The agentic enterprise This spring, a new artificial intelligence model quietly edged out some of the world’s best climate prediction systems. ### When Companies Begin to Think as Cognitive Enterprises Published: 2025-05-16 · LinkedIn: https://www.linkedin.com/pulse/when-companies-begin-think-cognitive-enterprises-dinand-tinholt-q0wac Theme: The agentic enterprise In late winter, a global logistics firm noticed something odd in its real-time dashboards. ### AI Agents on Autopilot: A Business Wake-Up Call Published: 2025-05-13 · LinkedIn: https://www.linkedin.com/pulse/ai-agents-autopilot-business-wake-up-call-dinand-tinholt-jwgpc Theme: The agentic enterprise For years, artificial intelligence has quietly shadowed the world of work. It assisted, suggested, and supported — playing the role of co-pilot. But something is changing. ### When Software Learns the Job Retail Finds Another Gear Published: 2025-05-12 · LinkedIn: https://www.linkedin.com/pulse/when-software-learns-job-retail-finds-another-gear-dinand-tinholt-xgfzc Theme: Retail & consumer goods Just imagine a store manager at a grocery chain juggling numerous urgent tasks: a customer wants to know why last night’s curb-side order is late, a district merchandiser is flagging a surprise run on energy drinks downtown, and… ### Why Tariffs Are a Data & AI Problem—And How Capgemini Is Solving It Published: 2025-05-07 · LinkedIn: https://www.linkedin.com/pulse/why-tariffs-data-ai-problemand-how-capgemini-solving-dinand-tinholt-cnysc Theme: The agentic enterprise In a year where AI seems to be injected into every boardroom conversation, one of the most impactful applications is happening at the border. ### What the AI Model Market Quietly Reveals About How to Build with It Published: 2025-05-05 · LinkedIn: https://www.linkedin.com/pulse/what-ai-model-market-quietly-reveals-how-build-dinand-tinholt-cbqvc Theme: Models & the AI market The conversation around AI models often focuses on performance. Speed. Accuracy. Cost. And while those factors matter, they don’t explain the full picture of how businesses are actually using large language models. ### Beyond Generative: The Rise of the Simulated, Self-Learning Enterprise Published: 2025-04-29 · LinkedIn: https://www.linkedin.com/pulse/beyond-generative-rise-simulated-self-learning-dinand-tinholt-zvcdf Theme: The agentic enterprise We are entering a profound shift — one that goes far beyond new technology. Generative AI lit the first spark, but what comes next will redefine how we understand intelligence itself. ### Rethinking Intelligence: Thriving in a World Where AI Outgrows Us Published: 2025-04-26 · LinkedIn: https://www.linkedin.com/pulse/rethinking-intelligence-thriving-world-where-ai-outgrows-tinholt-o5hic Theme: Work, economy & the future In a recent conversation that captured the attention of the technology world, Eric Schmidt, the former CEO of Google, shared a stark prediction: within the next year, the majority of programmers could be replaced by AI. ### From Prediction to Purpose: The Era of Experiential Intelligence Published: 2025-04-22 · LinkedIn: https://www.linkedin.com/pulse/from-prediction-purpose-era-experiential-intelligence-dinand-tinholt-hl11c Theme: The agentic enterprise How World Models and Reinforcement Learning Are Reshaping Business… ### The Rise of the Self-Reflexive Enterprise: Why the Next Wave of AI Isn’t About Doing More — But About Seeing Better Published: 2025-04-20 · LinkedIn: https://www.linkedin.com/pulse/rise-self-reflexive-enterprise-why-next-wave-ai-isnt-doing-tinholt-wgk8c Theme: The agentic enterprise On an average everyday morning, a thousand things are happening inside a multinational company. A frustrated engineer is rewriting the same process that’s already been documented elsewhere. ### The New Competitive Edge: Open-Weight AI Models and Their Impact on Businesses Published: 2025-04-16 · LinkedIn: https://www.linkedin.com/pulse/new-competitive-edge-open-weight-ai-models-impact-dinand-tinholt-pa8xc Theme: Models & the AI market The AI landscape is undergoing a seismic shift. In just the past month, OpenAI and other major players have signaled a new era—one where the power and potential of open-weight large language models (LLMs) are reshaping the… ### The Accelerating Reality of Artificial General Intelligence: Strategic Imperatives for Business Leaders Published: 2025-04-08 · LinkedIn: https://www.linkedin.com/pulse/accelerating-reality-artificial-general-intelligence-business-dinand-8bmgc Theme: Work, economy & the future Artificial General Intelligence (AGI), once confined to theoretical discussions and science fiction, is rapidly approaching practical realization. ### A New Solution for Trade Wars: How Enterprises Can Regain Control with Strategic AI Published: 2025-04-04 · LinkedIn: https://www.linkedin.com/pulse/new-solution-trade-wars-how-enterprises-can-regain-control-tinholt-fpljc Theme: The agentic enterprise As global trade becomes increasingly unpredictable, executives are finding themselves caught in a crossfire of shifting tariffs, regional restrictions, and disrupted supply chains. ### If AI tools were your office colleagues, who would they be? Published: 2025-04-02 · LinkedIn: https://www.linkedin.com/pulse/ai-tools-were-your-office-colleagues-who-would-dinand-tinholt-8qacc Theme: Ideas & culture Just for fun — and because we all need a breather from the firehose of AI news — here’s how I imagine… ### The Day the Machines Fooled Us Published: 2025-04-02 · LinkedIn: https://www.linkedin.com/pulse/day-machines-fooled-us-dinand-tinholt-datkc Theme: Ideas & culture What GPT-4.5’s Turing Test triumph means for business—and for being… ### The AI Retail & Consumer Goods Revolution: Five Stages of Transformation Published: 2025-03-31 · LinkedIn: https://www.linkedin.com/pulse/ai-retail-consumer-goods-revolution-five-stages-dinand-tinholt-hzcbf Theme: Retail & consumer goods The shop floor at Walmart's newest "store of the future" in Arkansas appears ordinary at first glance. Behind the scenes, however, an invisible intelligence silently orchestrates nearly every aspect of the operation. ### AI Doesn’t Want Your Job. It Wants Your Judgment. Published: 2025-03-31 · LinkedIn: https://www.linkedin.com/pulse/ai-doesnt-want-your-job-wants-judgment-dinand-tinholt-jgmic Theme: Work, economy & the future When earlier this week Bill Gates predicted that artificial intelligence would replace doctors and teachers within the next decade, it sounded hyperbolic—like a headline designed to provoke. ### When Intelligence Joins the Team: Rethinking Work in the Age of AI Published: 2025-03-25 · LinkedIn: https://www.linkedin.com/pulse/when-intelligence-joins-team-rethinking-work-age-ai-dinand-tinholt-suzic Theme: Work, economy & the future There’s a quiet revolution unfolding in how we work. And like many revolutions, it began not with a bang, but a study—one involving 776 professionals at Procter & Gamble and an AI model named GPT-4. ### The Rise of Small Giants: How Tiny AI Models and Massive Chips Are Reshaping Business Published: 2025-03-19 · LinkedIn: https://www.linkedin.com/pulse/rise-small-giants-how-tiny-ai-models-massive-chips-business-tinholt-oxysc Theme: Models & the AI market The story of artificial intelligence has long been about scale—bigger models, more parameters, and an insatiable demand for computing power. ### Antifragile — The Algorithm That Ate Uncertainty: How AI is Reinventing Supply Chain Resilience Published: 2025-03-16 · LinkedIn: https://www.linkedin.com/pulse/antifragile-algorithm-ate-uncertainty-how-ai-supply-chain-tinholt-ydfhc Theme: The agentic enterprise In March 2021, as the Ever Given — a container ship longer than the Empire State Building is tall — lodged itself into the banks of the Suez Canal, a chilling realization rippled through global markets: modern supply chains are… ### Pi, AI, and the Limits of Predictability Published: 2025-03-14 · LinkedIn: https://www.linkedin.com/pulse/pi-ai-limits-predictability-dinand-tinholt-5hrzc Theme: Ideas & culture Today is March 14th—Pi Day, a celebration of the most famous irrational number in history. Pi is a mathematical constant, an unbroken thread that runs through geometry, physics, and even the algorithms that drive AI. ### Retail’s AI Transformation: A Moment of Reinvention Published: 2025-03-10 · LinkedIn: https://www.linkedin.com/pulse/retails-ai-transformation-moment-reinvention-dinand-tinholt-lovnc Theme: Retail & consumer goods On a quiet Tuesday morning in early 2025, a fashion buyer at a global retail brand sat in front of her laptop, scrolling through trend reports, sifting through spreadsheets, and making a judgment call on next season’s colors. ### The Coming of AGI: Why Organizations Must Think Big and Act Now Published: 2025-03-04 · LinkedIn: https://www.linkedin.com/pulse/coming-agi-why-organizations-must-think-big-act-now-dinand-tinholt-i33zc Theme: Work, economy & the future Artificial General Intelligence (AGI) is no longer a distant dream. It is accelerating toward us much faster than previously expected. ### The Productivity Paradox: Why AI's Promise May Still Be Unrealized in Today's Economy Published: 2025-03-02 · LinkedIn: https://www.linkedin.com/pulse/productivity-paradox-why-ais-promise-may-still-todays-dinand-tinholt-hmwcc Theme: Work, economy & the future In a recent Bloomberg article, "AI Will Upend a Basic Assumption About How Companies Are Organized," ( https://www.bloomberg.com/news/articles/2025-02-28/how-ai-reasoning-models-will-change-companies-and-the-economy ) a… ### The Unseen Hand: How AI is Quietly Rewriting Retail’s Future Published: 2025-02-26 · LinkedIn: https://www.linkedin.com/pulse/unseen-hand-how-ai-quietly-rewriting-retails-future-dinand-tinholt-4phvc Theme: Retail & consumer goods A century ago, department stores revolutionized retail. Fifty years later, shopping malls did it again. Then came e-commerce, reshaping consumer behavior in ways no one could have imagined. ### Why Testing AI Like a Smartphone is a Mistake – And What Consumer & Retail Companies Should Do Instead Published: 2025-02-25 · LinkedIn: https://www.linkedin.com/pulse/why-testing-ai-like-smartphone-mistake-what-consumer-retail-tinholt-clwec Theme: Retail & consumer goods Imagine you’re evaluating a brand-new smartphone, but instead of exploring its camera, apps, or AI-driven features, you judge it solely by how well it makes phone calls. That would be absurd, right? ### The AI Boom is Bigger Than Satya Nadella Thinks—Here’s Why the Economic Impact Will Surpass Expectations Published: 2025-02-25 · LinkedIn: https://www.linkedin.com/pulse/ai-boom-bigger-than-satya-nadella-thinksheres-why-economic-tinholt-jr1lc Theme: Work, economy & the future Satya Nadella, CEO of Microsoft, has struck a cautious tone on the economic impact of artificial intelligence (https://futurism.com/microsoft-ceo-ai-generating-no-value). ### Reimagining the Consumer Products & Retail Market with GenAI: Four Futures, Five Years From Now Published: 2025-02-20 · LinkedIn: https://www.linkedin.com/pulse/reimagining-consumer-products-retail-market-genai-four-dinand-tinholt-uziyc Theme: Retail & consumer goods Five years from now, the consumer products and retail (CPR) industry will be unrecognizable. Not because of incremental improvements, but because of a seismic shift—one driven by the rise of Generative AI (GenAI) and Agentic AI . ### OpenAI’s Deep Research: The Next Leap in Work Transformation Published: 2025-02-11 · LinkedIn: https://www.linkedin.com/pulse/openais-deep-research-next-leap-work-transformation-dinand-tinholt-rm6cc Theme: Work, economy & the future Last week, OpenAI launched Deep Research, a new capability that pushes the boundaries of how businesses and professionals interact with knowledge. ### The Quiet Revolution: How GenAI is Reshaping Consumer Products Published: 2025-01-23 · LinkedIn: https://www.linkedin.com/pulse/quiet-revolution-how-genai-reshaping-consumer-products-dinand-tinholt-0grgc Theme: Retail & consumer goods The GenAI maturation is happening quietly, behind the scenes, in the world of consumer packaged goods (CPG), retail, and distribution (CPRD). ### Symbiotic Intelligence Theory (SIT): Rethinking Human Potential in the Age of AI Published: 2025-01-23 · Medium: https://medium.com/@tinholt/symbiotic-intelligence-theory-sit-rethinking-human-potential-in-the-age-of-ai-4664c8840652 Theme: Work, economy & the future A framework for thinking about human agency and the ways AI can extend our capacity to think together. ### AI Isn’t Just Replacing Jobs - It’s Rewriting Business Rules Published: 2025-01-15 · LinkedIn: https://www.linkedin.com/pulse/ai-isnt-just-replacing-jobs-its-rewriting-business-rules-tinholt-pfegc Theme: Work, economy & the future Imagine being the CEO of a $6.7B company and declaring that your business doesn’t need humans anymore. ### Key Trends from NRF 2025: Shaping the Future of Retail Published: 2025-01-14 · LinkedIn: https://www.linkedin.com/pulse/key-trends-from-nrf-2025-shaping-future-retail-dinand-tinholt-xq3oc Theme: Retail & consumer goods It's a wrap for 2025's NRF Retail Big Show. It's clear that 2025 is set to be a transformative year for the retail industry. ### We’ve been talking about major productivity gains from AI agents. Published: 2025 · Medium: https://medium.com/@tinholt/weve-been-talking-about-major-productivity-gains-from-ai-agents-796b2bd5a18c Theme: Work, economy & the future Their internal study of 132 engineers reveals something we need to start accounting for in our business cases: 27% of AI-assisted… ### The AI Iceberg Isn’t What You Think, and That’s Exactly Published: 2025 · Medium: https://medium.com/@tinholt/-d261ec1dc65d Theme: Work, economy & the future ### The Great Disconnect: Why the AI Revolution Is Missing from Economic Data Published: 2025 · Medium: https://medium.com/@tinholt/the-great-disconnect-why-the-ai-revolution-is-missing-from-economic-data-31187ee30377 Theme: Work, economy & the future ### From Language Models to World Models and Why the Conversation Just Got Real Published: 2025 · Medium: https://medium.com/@tinholt/from-language-models-to-world-models-and-why-the-conversation-just-got-real-3ebab3fd5aca Theme: Models & the AI market A few months ago I wrote a piece titled From Language Models to World Models: Building the AI Operating Layer for Business. ### Kimi K2 and the Quiet Shift in Enterprise AI Published: 2025 · Medium: https://medium.com/@tinholt/kimi-k2-and-the-quiet-shift-in-enterprise-ai-a6d143bd45a0 Theme: Models & the AI market Kimi K2 may turn out to be one of those quiet turning points that only later becomes obvious. ### Inside the Carnegie Mellon and Stanford Study: How AI Agents Work 88% Faster but Still Need Humans Published: 2025 · Medium: https://medium.com/@tinholt/inside-the-carnegie-mellon-and-stanford-study-how-ai-agents-work-88-faster-but-still-need-humans-e1b5cecb470b Theme: Work, economy & the future For years, the conversation in boardrooms and tech conferences has revolved around the same big question: will AI replace human workers? ### When’s the Last Time You Asked, ‘Did You Use PowerPoint?’ Published: 2025 · Medium: https://medium.com/@tinholt/-e7ead1fbc66c Theme: Ideas & culture Two days ago, Seth Godin wrote a piece I particularly enjoyed called ‘The Writer’s Room’. ### AI genius on demand? Published: 2025 · Medium: https://medium.com/@tinholt/ai-genius-on-demand-f5d0125bf291 Theme: Work, economy & the future I just read the paper ‘Genius on Demand: The Value of Transformative Artificial Intelligence’. ### Stop focusing only on prompts. Published: 2025 · Medium: https://medium.com/@tinholt/-39cbbb81c26c Theme: Data, context & memory Stop focusing only on prompts. Start engineering context. ### The Next AI Race: Why Smart Beats Big Published: 2025 · Medium: https://medium.com/@tinholt/-7645a9ca5d4e Theme: Models & the AI market The GPU arms race is ending. The efficiency war has begun. ### The Future of AI Isn’t Bigger Models, It’s Smaller Published: 2025 · Medium: https://medium.com/@tinholt/-b0d9b186767a Theme: Models & the AI market The center of gravity in AI is shifting from brute force generality to right-sized intelligence. ### Agentic AI: Autonomy Begins in the Unexpected Published: 2025 · Medium: https://medium.com/@tinholt/agentic-ai-autonomy-begins-in-the-unexpected-380af42e2709 Theme: The agentic enterprise I was listening to a presentation from my colleague Andreas Sjostrom this morning and a statement he made that stuck with me is that… ### Stop Waving Around ‘95% of GenAI Pilots Fail.’ It’s Not the Whole Story. Published: 2025 · Medium: https://medium.com/@tinholt/stop-waving-around-95-of-genai-pilots-fail-its-not-the-whole-story-dded82cad366 Theme: The agentic enterprise The MIT ‘State of AI in Business 2025’ report has been making waves with the headline that 95% of GenAI pilots fail. ### A New Chapter in Enterprise Intelligence: What GPT-5 Signals for Business Published: 2025 · Medium: https://medium.com/@tinholt/a-new-chapter-in-enterprise-intelligence-what-gpt-5-signals-for-business-8e3ce41f2694 Theme: Models & the AI market ### If your AI strategy still starts with a dashboard Published: 2025 · Medium: https://medium.com/@tinholt/-ac5303feb062 Theme: Data, context & memory ### The Future of AI: What Ethan Mollick Says Awaits Us in 2025 Published: 2024-12-11 · LinkedIn: https://www.linkedin.com/pulse/future-ai-what-ethan-mollick-says-awaits-us-2025-dinand-tinholt-oadzc Theme: Work, economy & the future We’re standing at the edge of a technological revolution, and artificial intelligence (AI) is at its heart. ### The Thousand-Day Countdown: Sam Altman’s Provocation and the Dawn of Super-intelligence Published: 2024-12-06 · LinkedIn: https://www.linkedin.com/pulse/thousand-day-countdown-sam-altmans-provocation-dawn-dinand-tinholt-t25fc Theme: Work, economy & the future Yesterday, Sam Altman of OpenAI dropped a subtle but seismic provocation at the DealBook Summit: super-intelligence is coming in a matter of "a few thousand days." That’s a little under a decade. ### OpenAI's ChatGPT Pro: Revolutionizing Business or Just Hype? Published: 2024-12-05 · LinkedIn: https://www.linkedin.com/pulse/openais-chatgpt-pro-revolutionizing-business-just-hype-dinand-tinholt-yagvc Theme: Models & the AI market OpenAI has launched ChatGPT Pro, a $200 per month enterprise-grade AI service. This move raises questions about the future of work, the accessibility of advanced AI, and OpenAI's strategic direction. ### AI ‘Buy With Pro’ Feature: A Game Changer for Retail Published: 2024-12-05 · LinkedIn: https://www.linkedin.com/pulse/ai-buy-pro-feature-game-changer-retail-dinand-tinholt-56pic Theme: Retail & consumer goods E-commerce is evolving at breakneck speed, and Perplexity’s innovative ‘Buy With Pro’ feature is leading the charge. ### From Chaos to Managed Complexity: Adopting an Ecosystem Mindset in Data & Analytics Published: 2024-11-01 · LinkedIn: https://www.linkedin.com/pulse/from-chaos-managed-complexity-adopting-ecosystem-mindset-tinholt-a3g6c Theme: Data, context & memory A case for treating modern data environments as evolving ecosystems rather than trying to force every source and decision into one rigid center. ### Decentralized Data Governance and Data Virtualization: Revolutionizing the Speed of Analytics Published: 2024-08-28 · LinkedIn: https://www.linkedin.com/pulse/decentralized-data-governance-virtualization-speed-dinand-tinholt-x8d8c Theme: Data, context & memory In the ever-evolving landscape of digital innovation, a quiet transformation is underway, one that could significantly alter the way businesses and organizations handle their most valuable asset: data. ### The Liminality of Data: A Glimpse Between the Thresholds Published: 2023-10-03 · LinkedIn: https://www.linkedin.com/pulse/liminality-data-glimpse-between-thresholds-dinand-tinholt Theme: Ideas & culture What the idea of liminality can teach data and AI leaders about working between an established order and one that has not yet taken shape. ### Generative: A dance of minds Published: 2023-05-18 · LinkedIn: https://www.linkedin.com/pulse/generative-dance-minds-dinand-tinholt Theme: Work, economy & the future In a study published by Eloundou et al. it is predicted that 80% of the U.S. workforce will have at least 10% of their work affected by #AI . This impact is increasingly impacting white collar jobs - the work of knowledge workers. ### Everyone's invited to the data party Published: 2023-05-15 · LinkedIn: https://www.linkedin.com/pulse/everyones-invited-data-party-dinand-tinholt Theme: Data, context & memory The world of data has long been confined to the dominion of data scientists, data engineers and IT experts, nestled snugly in their ivory towers of coding prowess and analytical acumen. ### Creating a data-powered culture Published: 2022-01-21 · LinkedIn: https://www.linkedin.com/pulse/creating-data-powered-culture-dinand-tinholt Theme: Data, context & memory A durable starting point for making data part of everyday decisions: connect technology, shared habits and organizational confidence. ### What Anthropic’s AI Advertising Stand and Legal Plugin Reveal About the Future of Enterprise AI Published: year not recorded · Medium: https://medium.com/@tinholt/what-anthropics-ai-advertising-stand-and-legal-plugin-reveal-about-the-future-of-enterprise-ai-1e97d8a1eb85 Theme: Models & the AI market Over the past few days, significant attention has been directed toward Anthropic’s latest announcements, including the release of a… ### What Novel Do You Want That Doesn’t Exist? Published: year not recorded · Medium: https://medium.com/@tinholt/what-novel-do-you-want-that-doesnt-exist-1f3c6647b27d Theme: Ideas & culture Over Christmas, I finished Upgrade by Blake Crouch and hit that familiar wall: the story ends, you want more, but there’s no sequel. ### The Right Tool for Each Step Published: year not recorded · Medium: https://medium.com/@tinholt/the-right-tool-for-each-step-fb444a35222e Theme: The agentic enterprise When your primary tool is a large language model, every business process starts looking like a conversation to generate. ### Agentic AI Needs Cognitive Goals Published: year not recorded · Medium: https://medium.com/@tinholt/agentic-ai-needs-cognitive-goals-021ea96bc787 Theme: The agentic enterprise ### Why Your AI Sounds Smart but Learns Nothing Published: year not recorded · Medium: https://medium.com/@tinholt/why-your-ai-sounds-smart-but-learns-nothing-20a33aeffb37 Theme: Data, context & memory People keep saying AI is getting smarter. Most of the time what they really mean is that it sounds better. It talks more smoothly. ### From Games to World Models: Why 2026 Will Be the Year AI Learns to Act Published: year not recorded · Medium: https://medium.com/@tinholt/from-games-to-world-models-why-2026-will-be-the-year-ai-learns-to-act-172378528219 Theme: Models & the AI market For the past two years, most leaders have experienced AI as a word machine: reading, writing, summarizing, and explaining. ### The Blueprint Is Showing Published: year not recorded · Medium: https://medium.com/@tinholt/the-blueprint-is-showing-b6b3ec4945b5 Theme: Work, economy & the future ### The Future of Intelligence: Insights from Demis Hassabis Published: year not recorded · Medium: https://medium.com/@tinholt/the-future-of-intelligence-insights-from-demis-hassabis-5089d3e59d0a Theme: Ideas & culture ### The Open Decision Head URL: https://dinand.com/research/decision-head/executive/ Section: Research · Decision Head · Executive briefing (23 September 2026) The Open Decision Head Dinand Tinholt, Capgemini, Head of AI Center of Excellence, Americas 23 September 2026 Headline results. The business question Every governed process has a moment where a case could go several ways and something has to pick one. In a supplier onboarding queue, an accounts-payable exception or a claims file, the person at that moment is asking whether we have what we need, whether to ask for something, whether someone should think harder about what's already here, or whether to close it. A wrong call there means a case closed early with an open obligation nobody caught, or hours of reasoning spent on a file that was missing one document all along. TypeSafe, a vendor, built a product called Jev around this kind of decision. You hand it a case and a list of yes/no or multiple-choice questions, and it answers each one with a calibrated probability, in 0.35 seconds and for $0.0000128 a question when we measured it. This brief looks at what happens when a company builds the same idea on its own hardware, and at what an operator gets and gives up by doing so. What a gate decides at each checkpoint. What the open decision head offers There are nine advantages, and each one follows from running the service yourself instead of sending a case to someone else's. Data stays on your network. The case and every answer about it stay on hardware the company already controls, documents included. A pilot can begin before legal has finished a data-processing agreement, since there's no vendor boundary for that agreement to cover. Calibration you own. The confidence number behind each answer is fitted on the company's own labelled cases, and an auditor can inspect it as a reliability table. When the company's documents change, the calibration gets refit. A hosted service's stated confidence stays as the vendor fitted it. Customizable in about a day. Adding a use case takes a new question file and a calibration run, and the integration and the contract stay as they are. With a hosted general model, each process goes through its own validation. The open head reaches a second process through the same integration, in about a day. Fine-tunable, with a measured payoff. Once enough of the company's own decisions have piled up, a smaller model can be distilled on those cases, faster and calibrated for that one process. A first version of this has run, and I report it later in this brief. It scored 0.701 against the full-sized retrained head's 0.728, short of the bar for replacing it, so it carries an experimental label. Audit per decision. Every answer carries a row with the exact question, the raw probability, the model that answered and how long it took. You can rebuild a decision from six months ago in full. Base model of choice. The same service runs on the model this project uses today (DeepSeek V4.1 Flash, served as one model spanning two GPUs on the fleet's own hardware). It also runs on a small model on a single workstation, or on a frontier model reached through the company's own cloud account, and the approach stays the same. Isolation checked on your own data. Jev's premise is that each question gets answered without seeing the others. The open version tests that directly on the company's data and keeps the result where a reviewer can check it. Cost from the power bill. Running cost comes off the electricity bill with no signup step, and capacity is whatever the company's own hardware provides. The model file is pinned by its own fingerprint, so nobody can quietly swap it underneath a company's own testing. Open source. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head , after two independent security reviews (23 September and 6 October 2026). A company's security team can read every line that touches its data before the first real case goes through. The hosted service keeps some advantages. Its model was trained specifically for this kind of typed decision, where the open head reads a general model one token at a time. It publishes a (self-reported) accuracy figure across its own evaluation set, and there's a vendor behind its support line. The experiment below puts a measured number on the gap between the two. What this looks like in practice (an illustrative smoke test, outside the benchmark). In the service's own validation run, a supplier's insurance certificate named the parent company instead of the subsidiary applying for onboarding. Asked whether the certificate met the requirement, the service said no, at 99.9 percent probability. Then it was asked what should happen next, with continue, request more information and escalate to a person as the options. It picked request with probability 0.91, correctly reading the gap as one the supplier could fix. On a three-level urgency scale it answered medium, but only just, with its probability split almost evenly across all three levels. That's the kind of case a company would want a person to glance at before anyone closes it on autopilot. Once the service had looked at the case once, all three answers came back in under three seconds. The fleet's governance log for 23 September 2026, wiki/governance/log-2026-09-23.md , records the satisfied and request readings. Why self-hosted matters to a Fortune 500 buyer Four reasons come up in nearly every procurement conversation about this kind of technology, and each maps to one of the properties above. Data residency and vendor risk review usually take longer than anything else before a pilot starts. A self-hosted service removes the reason for that review to exist. Owning the calibration fit gives a compliance function a confidence number it can defend, where a vendor's number has to be accepted as stated. A vendor waitlist, or a vendor closing enrollment altogether, is a live risk here. TypeSafe stopped issuing new API keys on the day this project started, and that's what triggered building the open version at all. Running on hardware the company already owns turns a per-decision invoice into whatever that hardware already costs to run. Speed versus accuracy: what matters here Measured on 24 September 2026, Jev, reached through OpenRouter, answered a single question in 0.35 seconds at a cost of $0.0000128 for 304 input tokens. The open head's fast mode was tested under real load on the fleet's own hardware, two DGX Sparks running four benchmark streams at once, and the results sit in PREREG-J6.md and results/design-review-J6-2026-09-24.md : measure value median response time 3.1 seconds p90 response time 4.8 seconds This run didn't bank a worst-case latency or a request count. Jev is the faster service. A deliberate mode that samples eight answers per decision, meant for harder cases, will cost roughly eight times a single fast call. Nobody has measured its speed yet. In most enterprise workflow gates the latency gap changes nothing. A gate fires once per checkpoint in a multi-step process. The J5 eval2 run logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case. The work around each gate, such as fetching a document from a supplier or routing a file to an expert, takes minutes to days. A three-second decision inside a step that already waits two days adds almost nothing to the wait, and a wrong call costs far more. Seconds, logarithmic scale. On 800 eval cases, the open head's zero-shot accuracy was 0.545 against a hand-written rule table's 0.575. A statistical check puts that gap of -0.029 somewhere between +0.001 and -0.060. Jev scored 0.540 on the same eval set, 0.560 on a development set and 0.536 on a 200-case set built with contradictory evidence, and its own gap against the rule table excludes zero on both sets. On that first run the rule table finished ahead of both the zero-shot open head and the hosted service, though the open head's gap sat inside the interval. On the contradictory-evidence set, every approach tested, the rule table included, wrongly called a case complete 21 to 28 percent of the time under the original scoring rule. A review afterward traced most of that to a defect in the scoring rule itself, because the test sometimes offers two equally valid documents and marks an answer wrong for citing the second one. Under a corrected scoring rule the rate falls to between 1.5 and 9.5 percent depending on the approach. A second round retrained the open head on the same task and tested it on 800 fresh cases. It scored 0.728 there against 0.566 for the rule table. The harder set showed the same pattern, with 0.739 against the rule table's 0.565, which put the head 0.199 ahead of Jev's own 0.536 on the 199 cases both scored. Every one of those gaps excludes zero, so this is the first result in the project where the open head beats both the rule table and the hosted service. On the fresh set, the retrained head's rate of wrongly calling a case complete first came in at 0.047. That's above the level the project set in advance as the trigger for redesigning a safeguard. When we traced why, the cause sat in the environment that builds the evidence pack. It was listing extra supporting documents it never checked, most of them decoys planted to catch a careless citation, so the error belonged to the pack builder. Once it checks every document the way the rest of the system already does, the rate drops to these figures: rules: 0.0088 retrained head: 0.0113 learned head: 0.0163 Every case that changed went from wrong to verified, and no other case moved. The pre-registered bar is met on the fixed reading and missed on the original one. We report both, because the fix was found after the bar had been missed. Share of cases closed with a claim the registered scorer marks wrong. For a checkpoint decision I'd rank a gate first on accuracy against the company's own cases and on keeping the data inside the company's network. Next come a model that retrains on a hundred of the company's own labelled cases with no vendor ticket involved and an audit row behind every call, with speed after those. Anyone choosing should weigh Jev's strengths too. It wins outright on latency and on needing no infrastructure, and it learned from a training set far larger than anything one company logs on its own. What the experiment showed A registered test, described in a companion technical paper, compared five ways of choosing the next step in a synthetic supplier-onboarding workflow. The first was a rule table and the second a local model choosing freely. Then came the same local model asked the identical typed questions the decision head asks, the open decision head itself, and TypeSafe's own Jev, reached through OpenRouter. All three local approaches ran on the same model, the fleet's own DeepSeek V4.1 Flash, so any difference between them comes down to how the question was asked. The test measured whether each approach can tell "we need a document" apart from "we need to think harder," and whether acting on that distinction finishes more cases correctly. The predictions were on the record before anyone opened a single evaluation case. In the first run the open decision head landed slightly below the rule table, where we had predicted it would land above the free-choosing model, and Jev landed below the rule table too. The second run tested a retrained head on a fresh set of cases, and that head beat both. The table below reports both runs. Each dot is one approach's difference in decision-class accuracy from the rule table on the same cases. The right-hand column gives the accuracy itself. Fine-tuning and retraining used the 600-case onboarding training split and the 100-case accounts-payable development split. Results. what J1, first run J5, retrained head, fresh cases J6, second process (AP) decision-class accuracy, rule table 0.575 (eval) 0.566 (eval2) / 0.565 (challenge) 0.648 (ap-eval) / 0.638 (ap-challenge) decision-class accuracy, open head 0.545 (eval), below the rule table 0.728 (eval2) / 0.739 (challenge), above the rule table 0.650 (ap-eval) / 0.659 (ap-challenge), level with the rule table decision-class accuracy, Jev (hosted) 0.540 (eval) / 0.536 (challenge), below the rule table not re-run on the fresh set 0.636 (ap-eval) / 0.622 (ap-challenge), at or below the rule table and the head decision-class accuracy, learned head not applicable 0.701 (eval2), retrained on the first process, experimental 0.522 (ap-eval) / 0.558 (ap-challenge) untrained on the new process; 0.663 / 0.671 once retrained on 100 of the new process's own labelled cases consistency, share of repeats that disagree 0.05% (Jev), 4.95% eval / 5.85% challenge (open head), while a second job shared the endpoint not re-measured on this arm re-measured on an isolated endpoint by a follow-on test: 0.05% at one call in flight, matching Jev exactly; rising to 1.95% at four calls in flight wrong-claim rate, retrained head not applicable 0.047 (eval2), fixed to 0.0113 once the evidence pack was corrected not read the same way; see below claim guard not applicable not needed once the evidence pack was corrected; two model-based attempts added nothing beyond that fix not tested on the new process decision-class accuracy, Laya (fine-tuned) not applicable 0.815 (challenge) / 0.824 (eval2, 400-case subsample), both above the judging head 0.554 (ap-challenge) / 0.539 (ap-eval), both below the rule table decision-class accuracy, CLM-8B (fine-tuned) not applicable 0.687 (challenge) / 0.641 (eval2 subsample), above the rule table, below the head on eval2 0.568 (ap-challenge) / 0.586 (ap-eval), below the rule table item-level calibration (ECE) not applicable judging head 0.0178 uncalibrated / 0.0143 refit on full eval2, 0.0313 uncalibrated on the 400 shared cases; Laya 0.0057 uncalibrated / 0.0054 refit on the same 400 cases not measured on this process Every accuracy gap above is a statistically real difference apart from the first run's open head against the rule table, where the interval runs from +0.001 to -0.060 and the result stays inconclusive. J1 also measured whether the head's stated confidence tracked its own correctness, and the first head's calibration was poor there. Its expected calibration error came in near 0.52, roughly six times the registered 0.08 bar. The item-level calibration reported elsewhere in this document belongs to a later head and is a separate figure, from H20. That head was already well calibrated before any refit. Repeating the same decision twenty times gave two different readings depending on how the test was run. In the first run a second job was sharing the endpoint. Jev changed its answer on 0.05 percent of repeats, while the open head changed on 4.95 to 5.85 percent across splits, so on that reading Jev looked like the steadier decider. A follow-up test re-ran the open head with nothing else on the endpoint, replaying the same points once at one call in flight and once at four calls in flight. At one call in flight the open head changed its answer on 0.05 percent of repeats on both the onboarding and accounts-payable processes, which matches Jev's own figure exactly. At four calls in flight it rose to 3.4 percent on onboarding and 1.95 percent on accounts payable. Those disagreements clustered on a single worker thread, the signature of calls being batched together, which places the earlier gap in the serving stack. The shipped service now defaults to one call in flight, which costs throughput and leaves accuracy where it was. Where all twenty repeats agree, more than half of that agreement lands on the wrong answer, and that holds for both approaches. Share of repeats that disagree with the majority answer, twenty repeats per decision point. A separate attempt tried to hedge the disagreement that remains by having the service abstain on the claims it is least sure of. The mechanically chosen band abstained on 1.8 percent of claims on one accounts-payable split and 0.7 percent on the other. At rates that low you can't separate the band's own effect from the ordinary noise of an independent rerun, so the test is reported as inconclusive. A cleaner version is already registered. It reads the band's effect from a single run's own trajectory instead of comparing two separate runs. Another test was meant to rank Jev's chosen actions against the rule table's by replaying every alternative from a saved checkpoint. Because the rule table finishes every arm's branches, the test favors the rule table by construction, so it's being rebuilt and won't be reported as a ranking. The evaluation set (eight hundred held-out cases, plus a two-hundred-case harder set with shifted formats and a policy change) was built and frozen before either run. An independent review checked the protocol three times before it was allowed to run, and once more before the retrained head's run. Afterward, a party who did not build the scoring code rescored both runs from the raw case files. Carrying it to a second process Everything above ran on one synthetic process, supplier onboarding. A follow-on test moved the same approaches, Jev included, onto a second process built the same way: accounts-payable exception routing. On the new process's evaluation split the judgement head scored 0.650 and the rule table 0.648, level with each other, and the head's lead from the first process was gone. A learned head trained on the first process scored 0.522 there when pointed at the second one with no retraining. Trained instead on a hundred of the new process's own labelled cases, 271 decisions in all, it matched the judgement head. Budget for roughly a hundred labelled cases from your own process before you trust a learned head on it. A follow-up check showed that the match holds without the learned head seeing the judgement head's own answers while training or while running, a question this brief had left open. A retrained version had those answers stripped out of its features and was served with no call to the judgement head at all, and it still matched. It scored 0.660 against the stacked version's 0.663 on one split, and 0.670 against 0.671 on the other, with both gaps well inside noise. On the larger split it also wrongly called a case done less often than every other approach. Jev, tested on the same second process, came in at or below both the rule table and the judgement head. Its caution showed up as fewer wrong claims of a case being done and more unnecessary hand-offs than either, with accuracy at or below the other two. Share of cases, Jev minus the judging head, with both sides scored on validated evidence packs. Left of zero means Jev does it less often. This second process was built the same way as the first and shares its mix of case types. Whether any of this holds on a process built differently is still open, and that test comes next. What it costs The evaluation runs, on hardware the company already owns, cost nothing beyond electricity. A narrow, capped budget was set aside for two jobs that need a stronger outside model: consulting an expert model on the hardest cases, and a blind check of results by a model that never saw how any answer was produced. That phase hasn't run yet, and the sixty-dollar ceiling stays reserved until it does. Jev didn't need a reopened TypeSafe account. We reached it through OpenRouter at the $0.0000128 per question measured on 24 September, and that cost is already folded into the results reported above. The limits of the claim The claim here stops short of saying a company's own hardware beats a well-funded vendor's model at this task in general. It rests on a controlled comparison whose results point both ways. The first version of the open head didn't beat the rule table, and a retrained version did, by a wide margin, on two different test sets. On production readiness the claim is narrower than the accuracy numbers alone would suggest. The retrained head's wrong-claim rate first came in above the line this project set in advance. Three guards were tried for that failure mode, two of them model-based, and both of those failed their own test. A vote across several readings separated a wrong claim from a right one at an AUC of 0.56. A version that read the cited document itself did no better than a plain deterministic check, at an AUC of 0.52. None of the three is in use. The wrong-claim rate came down once we corrected the evidence pack, which the environment had been filling with supporting documents it never checked, and the corrected rate sits close to the rule table's own. We found that fix after the bar was first missed, by tracing why, and the results above say so. The claim does cover direction, now backed by a measured result as well as a prediction. A company can reproduce the pattern behind a typed-question decision service without a vendor and tune it on its own cases. It can measure the result in the open and run it on its own terms, provided the process around the model checks its own evidence too. What a company will be able to do with the code The service is a few hundred lines of Python with no third-party library required, and it runs in front of any model that exposes token-level probabilities, including one a company already has behind its own firewall. Once it is released, a team can point it at their own ollama or vLLM endpoint and write a handful of yes/no or multiple-choice questions about a case they already handle, and get calibrated answers, each with an audit trail. It is public on GitHub under Apache-2.0 ( https://github.com/dtinholt/decision-head ), after two independent security reviews, so a company's own team can read every line before trusting it with real cases. The open-weights step: a first result The first version of this step has now run. Alongside the retrained decision head above, a separately fine-tuned, smaller head was trained on the project's own labelled cases and registered as its own arm. It scored 0.701 on the fresh evaluation set, close to the retrained head's 0.728 without matching it. That gap doesn't clear this project's own bar for calling the two equivalent, so the fine-tuned version ships with an experimental label and stays out of the replacement role. What the open alternatives showed Two open encoders released by vendors, Laya and CLM-8B, appeared within days of this project's own head. Each answers a decision in one forward pass without reading the evidence behind a case. We ran both on this project's own benchmark instead of taking either vendor's numbers at face value. Fine-tuned on 1,567 of our own labelled decisions, drawn from 363 onboarding cases, a small open encoder scored 0.815 on that process's harder set, above our judging head and well above the rule table's 0.565. The same encoder fell below the rule table it was supposed to beat when it had only 201 labelled decisions from 58 cases of a second process to learn from. Untuned, both encoders matched the rules and went no further. For this encoder the size of its training set decided the result, since the smaller accounts-payable set left it with a habit that cost accuracy on that process. Decision-class accuracy on the harder set of each process, 200 cases each. We also checked calibration, since one vendor advertises it as a strength. The judging head's own probabilities were already well calibrated before this test ran. On the 400 shared cases the open encoder's calibration error was 0.0057 against the judging head's 0.0313, both small. The judging head reads the cited evidence and keeps the ledger, with an audit row behind each call, and an encoder does none of it. The judging head stays the safe default. A learned mode gets promoted only after it passes a held-out check on your own data, whatever its vendor reports. ### The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control URL: https://dinand.com/research/decision-head/paper/ Section: Research · Decision Head · Paper (23 September 2026) The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control Dinand Tinholt Capgemini, Head of AI Center of Excellence, Americas 23 September 2026 This paper describes a registered experiment and the service it tests. Two rounds of it have run, and an independent party has rescored both from the raw case files. J1 was the original four-way comparison. J5, a second registration, retrained the head and tested it on a fresh split. Each number below comes from those two verified rounds, states a prediction and says so, or cites someone else's published work by name. Section 6 sets the measurements against the five pre-registered hypotheses and explains why the claim guard built for the retrained head's wrong-claim rate gave way to a fix in the evidence bookkeeping. Abstract Much enterprise AI work inside a governed process consists of decisions made at checkpoints. At each one the system has to pick among retrieving a document, asking a supplier for one, reasoning over what it already has, escalating to a stronger model, closing the case, or handing it to a person. TypeSafe's Jev is built for this kind of question. It receives a case and a state once, answers a set of typed questions about that state in parallel and in isolation, and returns a calibrated probability for each. Two outside studies have measured parts of this claim. One checked variance on five fixed responses; the other was a post-hoc audit that ranked Jev against a frontier model at flagging wrong actions already taken. Both scored answers to fixed or finished cases. This paper registers a test of the next-step choice made while a case is still open, with the model's stated probabilities scored against its own record. It also asks if the pattern needs a hosted vendor at all. The vendor closed enrollment on the day the test was designed, so the paper also builds and evaluates a self-hosted equivalent. This decision head reproduces Jev's request and response shape over any local model that exposes token-level probabilities. It makes one single-token call per question, calibrates on the operator's own cases and writes an audit row for every decision. We describe the mechanism, the experiment (six actions, four decision classes, five controllers, an obligation ledger that catches silent omission, and counterfactual replay from saved checkpoints), and the five pre-registered hypotheses with numeric predictions. On the first registered split the self-hosted head failed to beat a hand-written rule table, and so did the hosted Jev service measured afterward on the same split. A second registration retrained the head and tested it on a fresh split, where it beat the rule table by 0.161 to 0.171 decision-class accuracy and beat Jev by 0.199 on the 199 cases both scored, all three intervals excluding zero. Its wrong-claim rate on that split sat above the threshold this project set in advance to trigger a guard redesign. The redesign traced the defect to the workflow engine's own evidence bookkeeping. Once the pack builder validates every supporting document before a claim leaves, the retrained head claims wrongly at close to the rule table's own rate, and no model-based guard added anything beyond that check. We found the fix after the threshold was first missed and report it that way. A third registration tested transfer to a second synthetic process built with the same generator family. There the rule table matched the judging head and a zero-shot transfer of the learned head scored 0.522 against the judging head's 0.650; retraining on a hundred labelled cases from the new process recovered parity. Jev, the hosted reference, handed off more cases on the new process and claimed fewer wrongly, with accuracy at or below the rule table and the judging head. A party who did not build the scorer rescored all three results from the raw case files. A fourth registration ran two vendor-released open alternatives, Laya and CLM-8B, on the same benchmark. Fine-tuned on 1,567 labelled decisions from 363 onboarding cases, one encoder beat the retrained head by 0.07 to 0.08 and the rule table by 0.25 to 0.28. The same recipe fine-tuned on only 201 labelled decisions from 58 cases of the second process fell 0.08 to 0.11 below the rule table. Both fine-tunes claimed a case complete wrongly less often than the judging head, and one encoder's item-level calibration matched the head's own. Across the encoder's own two training sets the larger set won, and the learned mode should be checked on held-out data before it is promoted. 1. Introduction A governed process depends on a small choice it makes over and over. Wherever the process could go one of several ways, something has to pick the way. In supplier onboarding or an accounts-payable exception queue, that choice seldom involves writing anything. The controller has to judge whether the evidence on file is sufficient, whether the case needs a document nobody has yet, whether interpreting what is on file needs a stronger model, or whether the case is done. Each wrong choice fails in its own way. Asking for more evidence when the answer was already on file wastes a cycle. Reasoning over a case that is missing a document produces a confident wrong answer. A compliance function worries most about a case declared complete while an obligation is still open, because nothing downstream catches it until an audit does. TypeSafe built a product around a narrower, more testable form of this problem. Jev takes a state and a set of typed questions, each phrased as a yes/no ("Noul"), a choice among named options, or a score. For each question it returns a probability distribution, computed in parallel and in isolation, at $0.042 per million input tokens with output free (docs.typesafe.ai, quickstart and primitives pages; typesafe.ai/blog/introducing-system-one-models-and-jev). The vendor's own evaluations page reports 67.8 percent average accuracy at $0.0004 per case and 0.4 seconds, self-reported, with a reference-model bias the vendor itself acknowledges (evals.typesafe.ai). Two outside parties have since tested parts of the claim independently. A LangChain post dated 20 September 2026 ran five fixed weather-agent responses through Jev one hundred times each and found the oracle answer reproduced on all five hundred calls, with per-case variance of 1.5e-5, at $0.00035 per call and 0.44 seconds. That run measured variance on five items whose answers were known in advance. Thordur Arnason's LinkedIn study, dated 23 September 2026, took completed agent trajectories on τ²-bench (320 conversations, 470 recorded database changes, 22 of them wrong) and asked whether Jev could rank the wrong changes for a human reviewer better than a frontier model. Jev scored 0.74 against Haiku 4.5's 0.77 (0.5 is chance) while running about four times faster at about four percent of the cost, and code-plus-narrow-questions doubled the catch rate inside the top ten percent of a review budget. He also ran the open-weight models Laya and Nemotron 3.5 Lightning on a single DGX Spark, and both ranked the same residual errors at about chance. Neither study tested the question a governed process needs answered most. Both reviewed actions after the fact. A controller has to decide what should happen next before any action is taken, while the case is still open and a wrong choice has time to compound. Thordur named the gap in his own write-up. In 31 of his 53 failed conversations the agent made no wrong database change at all; it failed by doing too little. A reviewer that only ranks recorded actions has nothing to rank in such a case. He also left unmeasured whether Jev's stated confidence tracks its own correctness. His stated next step is local fine-tuning, which raises a separate question this paper answers first, namely whether the value sits in the hosted model or in the shape of asking. If a local model given identical typed questions in an identical format performs close to Jev, the pattern carries the value and can run anywhere a model with token probabilities can be reached. If the local model falls well short, the hosted model is doing something local models cannot yet do. We register a prospective, next-step decision test with four things neither prior study measured: whether a controller can tell "the case needs a document that is not on file" from "the case needs interpretation of a document already on file," whether acting on that distinction completes more cases correctly at lower cost, whether the controller's stated probability of being right tracks being right, and whether an open, self-hosted model asked the same typed questions gets most of the way to a hosted one. TypeSafe closed enrollment to new API keys on the day this protocol was written, so to answer the fourth question we built a small service, described in Section 3, that answers Jev-shaped questions over a model we run ourselves. 2. Related work Jev and the System One family. TypeSafe describes Jev as a "System One" model, a fast, narrow model that answers typed questions, set against slower models built for open reasoning (typesafe.ai/blog/introducing-system-one-models-and-jev). The vendor documents nine known failure modes for the current version, jev-1.13.0 , including literal reading, arithmetic, dates, indirection, large irrelevant state, adversarial content, contradictory criteria, structural invariants, and generation (docs.typesafe.ai, model-jaggedness/jev-1.13). Few vendors publish where their own product breaks. This documentation shaped several design choices in Section 5, and it is why code handles arithmetic and date logic everywhere in this experiment. Routing. A separate line of work treats model choice itself as a decision problem, in which a query goes to a cheap model or an expensive one depending on predicted difficulty. RouteLLM and similar frameworks make this choice with a trained router placed in front of the models it routes between. The decision head in this paper is related and narrower. It answers a small set of fixed, typed questions about a state in order to choose the next action in a workflow, and several of those actions, such as retrieving a document or closing a case, involve no model call. Value of information and selective prediction. The INFO-versus-REASON distinction in this paper comes from decision theory. Before spending reasoning effort, the controller asks whether the missing piece is a fact that could be fetched or an interpretation that has to be made. Selective prediction lets a system abstain when its own confidence is low. The calibration literature measures whether a stated probability tracks empirical correctness. Both bear on whether a typed-question controller's probabilities are usable for anything beyond picking a top choice. Expected calibration error, the metric H3 is scored on, comes from the calibration literature. τ²-bench and post-hoc review. Thordur Arnason's study, summarized above, is the closest existing measurement of Jev's practical value and the direct predecessor to this registration. He ranked a finished trajectory after the fact, and this paper's design is built around a prospective next-step decision on an open case. LLM-as-judge consistency. Arm D in this design uses a frontier model as an expert and, in phase J3, as a blind adjudicator over a holdout sample. LLM judges are known to give inconsistent answers across repeated calls, orderings and phrasing, so the adjudication sample is drawn once from a fixed seed and scored blind, with no reruns in search of a preferred answer. Open single-pass alternatives. Two open decision models appeared within days of each other in late September 2026, described here as their vendors describe them in syn-open-decision-models-2026-09-28 , dated 28 September 2026. Laya, from Convai Innovations, released 18 to 21 September 2026 under Apache 2.0, is a 421-million-parameter ModernBERT-large encoder that answers typed choice, score, and yes/no primitives in a single forward pass, with a per-question-type temperature refit the vendor reports moving expected calibration error from 0.21 to 0.47 raw down to 0.08 to 0.11. CLM-8B, from Stanford and NVIDIA, released 23 September 2026 under Apache 2.0, is a frozen Qwen3-8B encoder behind two small contrastive heads that rank a given set of candidate actions by cosine similarity, and by the vendor's own account it has no calibration step. Both are single-pass encoders over typed primitives, and neither reads the evidence documents behind a case. Until now every published comparison between either model and Jev has come from a vendor or the press. Section 6.7 registers and reports the first run of both models on this project's own pre-registered benchmark. 3. The decision head Mechanism. A decision head takes one state and a set of typed questions about that state and returns one probability distribution per question. Three question types are supported: a Noul , a yes/no probability; a Choice , a probability over a small named set of options; and a Score , a probability over an ordered set of levels. The state is rendered once into a shared prefix (a system turn plus a user turn holding the state text), and every question is then asked as its own separate final turn appended to that same prefix, in its own model call. No question's prompt contains another question's text. That arrangement is what this implementation means by "evaluated in parallel and in isolation." The same arrangement lets a serving engine's prefix cache compute the shared prefix once, so the second and later questions in a batch cost less than the first. Figure 1. From state to audited answer. Each call asks the model for exactly one token, at temperature zero with logprobs turned on. The service reads the response's top-token log-probabilities and normalizes them over the tokens the question accepts ("Yes" and "No" for a Noul; a letter per option for a Choice or Score). Probability mass that fell on any other token is reported as unparsed mass. If none of the accepted tokens appear at all, the answer degrades to a flagged uniform distribution, so the service never fabricates a number. The read is mechanical. It takes a distribution the model already computes at that position, with no second model call and no sampling. The service tests isolation by answering the same question set twice. The first pass uses the normal isolated form. The second adds every other question's instructions to the shared prefix, so the model sees the full list while still being asked for one token at a time. The two runs are compared per question, reporting the raw probability drift and whether the winning answer changed. A client can therefore check a vendor's isolation claim as a number on their own state. Calibration applies temperature scaling to the logit of a probability: p' = sigmoid(logit(p) / T) , with the temperature fitted per question type on a labelled development set and applied at answer time. For a Noul question this is a direct transform of the single yes-probability. For a Choice or Score question, where more than one probability exists per question, the same transform is applied to every option and the result is renormalized to sum to one; this degenerates correctly to the Noul case when there happen to be exactly two options. Confidence for a Choice or Score answer is reported as the top probability minus the second, computed after calibration so the reported confidence always matches the reported probabilities. Every decision produces an audit row holding the question asked, the raw and calibrated probabilities, the winning answer, latency, and a hash of the exact request. The log holds only what was sent to the model, and secrets stay out of every log line. Where an API key is needed, the service reads it from a named environment variable at call time, and the key appears nowhere in code, configuration or output. A worked example, for illustration only. The service's own smoke test ( tools/decision-head , nine calls against qwen3.8:27b , the model used for this validation run) puts a real state through all three question types at once. State: "Acme Logistics BV, a subsidiary of Acme Holding NV, submitted a certificate of insurance dated 2026-01-15 naming Acme Holding NV as the insured party rather than the subsidiary. No endorsement schedule or subsidiary rider was included." A Noul asking whether the evidence on file satisfies the insurance requirement for Acme Logistics BV answered 0.001, so the model read the entity mismatch and said no with near certainty. A Choice over the next step ( CONTINUE , REQUEST , HANDOFF ) answered REQUEST at 0.910 against 0.007 for CONTINUE and 0.083 for HANDOFF , a margin of 0.827 over the next option. The certificate names the wrong entity, and the parent-subsidiary relationship makes this a gap the supplier can likely close. A three-level urgency Score answered medium by a small margin, at 0.188 low, 0.412 medium, 0.400 high, a margin of 0.012. The near tie shows the model uncertain between two adjacent levels, which is what a calibration step should surface. The three questions took 9.4 seconds on the first call and 2.7 seconds once the shared prefix was warm, and the isolation check (the alone-versus-together comparison above) changed none of the three decisions. qwen3.8:27b was the model available for this validation run and for the earlier single-token probe that first showed a logprobs read was possible; Section 5 describes the different, larger model the registered experiment runs its worker and arm C′ against. The satisfied and REQUEST readings above are recorded in the fleet's governance log for 23 September 2026, wiki/governance/log-2026-09-23.md . The API shape. The service exposes POST /v1/systemone , taking a state (string or object), a model field (read but not routed on: a client written against the hosted Jev API can point at this server by changing only the base URL), and a map of typed questions, and returning a matching map of answers, a usage block, and the model id the server was started with. A GET /health endpoint reports backend and model; GET /v1/models reports the real model id and any aliases a caller might send. The request's model field is advisory so that a hosted call can be replaced by changing one URL, with no code rewritten. Backends. Two real backends are implemented, sharing one interface ( token_distribution(prefix, question, allowed_tokens) -> distribution ). Ollama is called on its native /api/chat route only, with think:false . On the same server the OpenAI-compatible /v1 route passes through a thinking wrapper, and on the fleet's own host it returned the wrapper's filler token in place of an answer, so the service calls only the native route. An OpenAI-compatible backend covers vLLM-style servers and turns off template-level thinking through chat_template_kwargs.enable_thinking:false . Arm C′ calls this backend in the registered experiment, against the fleet's own dual-Spark endpoint (Section 5). A third backend generates deterministic fake distributions for tests and needs no model at all. The whole service uses the Python standard library with no third-party dependency, so it runs unmodified on a DGX Spark, on a workstation GPU, or on a small server in front of a hosted gateway. 4. What a self-hosted decision head adds and what it costs The case for a head built on the operator's own model rests on nine properties, each traceable to a design choice in Section 3, and it makes no performance claim against a hosted model. Data stays on the network. The state, the documents behind it, and every answer stay on hardware the operator controls, because the service and the model it calls both run there. A pilot can start before the data-processing agreement and the vendor security questionnaire are finished, since no vendor boundary exists for that paperwork to describe. Calibration the operator owns. The temperature-scaling fit in Section 3 runs on the operator's own labelled cases and produces a reliability table the operator can hand to their own auditor. A hosted service sets its stated confidence through its own process on its own cases, and the operator cannot refit it when their documents change. Any question set, any process. A new use case needs a new question file and a calibration run against a labelled sample. The isolation and audit mechanisms in Section 3 apply to any typed question, so one head can be pointed at a second process in about a day. Fine-tuning, with a measured transfer cost. Once enough operator-specific decisions have accumulated, a smaller model can be distilled on exactly those cases. That step is registered separately as phase J4, with its own pre-registration and adversarial review; a fine-tuned learned head, arm L, first ran in J5 (Section 6.3). Section 6.6 measures what a change of process costs, where a hundred labelled cases from the new process recovered the parity zero-shot transfer could not reach. Audit per decision. Every answer carries a row with the raw distribution, the exact model id, the question-set version, and the latency (Section 3). Someone reviewing a decision six months later can rebuild from that row alone what the model saw and returned, along with what a downstream system did with it. Choice of base model. The same service runs unmodified over the model this experiment now uses (DeepSeek V4.1 Flash, served as one model across two GPUs, Section 5), over a small model on one workstation, or over a frontier model behind the operator's own gateway. Only the backend configuration in Section 3 changes. Measured isolation. Section 3's isolation check answers a question set twice, once in isolation and once with every other question visible, and reports where the two disagree. The operator can produce that number on their own cases. Cost from the power bill. The service has no per-token charge and no signup queue, and it costs whatever the hardware already costs to run. The model file is pinned by hash, so no vendor can repoint an alias without notice. Open source. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head , release v0.1.0, after two independent security reviews (23 September and 6 October 2026), and an operator's own reviewers can read every line that touches their data before a single case is sent through it. All of this has a price, and a hosted, purpose-built model keeps advantages a self-hosted head does not inherit. jev-1.13.0 was trained for typed decisions. The decision head reads one token at a time from a general model never trained for that framing, and Section 5's H1 measures how much that gap costs. TypeSafe publishes a self-reported accuracy figure across four workflows on its own evaluation set. The decision head in this paper had no published track record beyond the pre-registered numbers its results section fills in. A hosted service also comes with a vendor SLA and a support line, while a self-hosted one gets whatever support the operator's own team is prepared to run. This paper takes no side on these differences. Section 5's experiment measures those that synthetic cases can measure today. What an enterprise gate should optimise for For most enterprise workflow gates the question that matters is whether the gate chose right, and the measurements below put speed and correctness side by side. Jev, reached through OpenRouter, answered a one-question probe in 0.35 seconds. The call cost $0.0000128 for 304 input tokens. Both measured 24 September 2026. The open head's fast mode, DeepSeek V4.1 Flash served across two DGX Sparks with four benchmark streams already contending for the same hardware, answered a yes/no question at the following latency. These figures come from PREREG-J6.md and results/design-review-J6-2026-09-24.md . No maximum latency or request count was banked for this run. measure value median latency 3.1 s p90 latency 4.8 s Jev answered faster on both of those numbers. Deliberate mode, which samples eight answers per decision for harder cases, is expected to cost roughly eight times a single fast call; its own latency and accuracy have not been measured yet. The latency numbers leave out the job a workflow gate does. The J5 eval2 run logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case before a case closes ( results/verification-J5-2026-09-26.md ). The steps a checkpoint gates, such as retrieving a document or routing a case to an expert, run minutes to days. A person waiting on a two-day retrieval step loses almost nothing to a three-second decision inside it. Correctness is where the arms differ. On the first registered split, neither the zero-shot open head nor the hosted Jev service beat the hand-written rule table; Section 6.1 reports the full comparison. On that same split's hardest cases, every arm appeared to claim a case complete wrongly in 21 to 28 percent of cases. Most of that traced to one scorer defect, and Section 6.4 reports the flawed reading beside the corrected one. A second registration retrained the open head and tested it on a fresh split, where it beat the rule table and Jev with every interval excluding zero (Section 6.3). Its wrong-claim rate sat above this project's own pre-registered trigger for a guard redesign, and the cause turned out to lie in the workflow engine's own evidence bookkeeping (Section 6.4). This paper reports the accuracy gain and the wrong-claim problem together, since either one alone would misstate what the project has shown. The head was designed with latency as a secondary concern. A self-hosted gate offers accuracy measured on the company's own process, and a live workflow's state stays inside the company's network. The model can be retrained on a hundred of the company's own labelled cases without a GPU. Every decision has an audit row behind it, and changing behaviour means relabelling data where a hosted service would need a vendor ticket. A hosted service answers in 0.35 seconds with no infrastructure to run, and it is tuned on far more data than any one company will log on its own. Those advantages are real, and they make Jev a serious reference for this project to measure against. For the gate this paper describes, adoption should turn first on accuracy and then on who owns the data and what it costs to change behaviour. 5. The workflow-completion experiment Task and environment. The controller under test sits inside a synthetic supplier-onboarding workflow. A local AI worker has already extracted findings from a case; the controller's job is to choose the next of six approved actions: CONTINUE , RETRIEVE(category) , REQUEST(item_id) , CONSULT_EXPERT , HANDOFF(item_id, issue) , or CLAIM_COMPLETE(pack) . Each case carries a policy version, an application, a set of requirements, documents already on file, a hidden internal store retrievable by category, what a supplier could additionally supply if asked, and a hidden key describing, for each open requirement, which of four resolution classes it belongs to. Every controller gets the same fixed budget per case of twelve steps, two expert calls and three requests. The four-way class. Every open item in a case needs exactly one of four things. Some items lack evidence; RETRIEVE resolves them if the evidence exists internally, and REQUEST if only the supplier can provide it. Other items need interpretation of evidence already on file, handled by CONTINUE when the local worker can make that judgment and by CONSULT_EXPERT when the case needs a stronger model. When every obligation is closed, the case needs CLAIM_COMPLETE . An item that no one can supply goes by HANDOFF to a person, with the issue named. The design names the two errors it is built to catch. info_as_reason spends reasoning effort on a case that was missing a document. reason_as_info asks for a document when interpreting the existing evidence would have covered the requirement. Figure 2. What a gate decides at each checkpoint. Arms. Five controllers are compared. Arm A is a rule table over the worker's findings and the obligation ledger, calling no model. Arm B is the worker's own model choosing freely among the six actions from a single JSON reply. Arm B′ asks that same model the identical typed questions the decision head asks, with probabilities estimated from eight samples at temperature one where the endpoint does not expose logprobs directly, mapped to an action through the same fixed decision table used downstream. Arm C′, the decision head described in Section 3, asks the same questions over the fleet's standing endpoint, one single-token call per question. The worker behind arms A, B, B′, and C′ is DeepSeek V4.1 Flash, a three-bit-per-weight EXL3 quantization served as one model across both of the fleet's DGX Sparks, reached over the OpenAI-compatible route described in Section 3. The original registration planned qwen3.8:27b as the worker, and PREREG-J1 addendum 4 records the change. The 27-billion-parameter model's role in this project ends at the single-token probe that first showed a logprobs read was possible and at the service's own validation smoke (Section 3); J1 measures a different model. Arm C, the hosted Jev service, ran through OpenRouter from J2 onward. A fifth arm, D, hands the entire workflow to a Bedrock model with no typed-question structure, as a cloud-first reference point with no causal role, and runs only in the credits-funded phase described below. A selftest asserts that B′ and C′ share byte-identical state compaction and question text, so any gap between them comes from the model. The obligation ledger and counterfactual replay. At the start of each case, one obligation is opened per requirement, closed only by a finding of met . A CLAIM_COMPLETE issued while any obligation is still open is recorded as a premature claim the instant it happens, whether or not a downstream guard later catches it. The ledger exists for the failure mode Thordur Arnason's study could not see, an agent that makes no wrong move and still does too little. Here that agent leaves an open obligation with no action taken against it, and the ledger records the gap even if no wrong action ever occurred. Every checkpoint can also be replayed. The environment snapshots its state before a decision and can restore it, so each of the six legal actions can be tried from the same starting point and scored against the same key. The resulting regret figure (best available outcome minus the outcome of the chosen action) separates the cost of picking wrong from noise in how the episode happened to unfold afterward. Metrics. The primary metric is four-way decision-class accuracy; exact-action accuracy is reported alongside as secondary. info_as_reason and reason_as_info are reported as rates on the checkpoints where each applies. Per case: verified completion (claimed, pack matches the key, every obligation closed), premature-claim rate, unnecessary handoffs on cases that were resolvable, correct blocks, expert calls, requests, steps, cost, and latency. For controllers that report a probability, calibration is scored with Brier score, expected calibration error over ten bins, a reliability table, and a selective-prediction curve showing decision error against coverage as the controller abstains below rising confidence thresholds. Consistency is measured on one hundred fixed checkpoints replayed twenty times each, as per-decision agreement and, separately, as the share of decisions that are consistently wrong. Hypotheses and predictions. Five hypotheses are pre-registered, each with a numeric prediction stated before any eval-split data exists. H1, discrimination. Arm C′ classifies the next-step decision more accurately than arm A and than arm B′, tested on the four-way class with a paired, case-clustered bootstrap confidence interval on the 800-case eval split. Predicted eval-split class accuracy: A, 0.62 to 0.74, point estimate 0.68; B, 0.50 to 0.66, point estimate 0.58; B′, 0.60 to 0.74, point estimate 0.67; C′, predicted between B′ and the hosted arm on every row, class accuracy 0.64 to 0.76, point estimate 0.70 (the hosted arm C, where it runs, is predicted at 0.68 to 0.82, point estimate 0.75). H2, completion at cost. Acting on C′'s choice yields more verified completions and fewer premature claims per dollar than arm A, at equal or fewer expert calls. Predicted: premature-claim rate 0.06 for A and roughly the same or better for C′; verified completion 0.55 for A against a predicted improvement for C′; expert calls per case roughly 0.9 for A against a predicted reduction for C′. H3, calibration. C′'s probabilities predict its own correctness (expected calibration error at or below 0.08 after temperature scaling fitted on the development split and reported on eval) in a way the local model's own sampled probabilities in arm B′ do not (expected calibration error at or above 0.15). Before temperature scaling, C′'s expected calibration error is predicted between 0.08 and 0.15. One structural limit is registered in advance. A correctly timed CLAIM_COMPLETE closes every obligation and therefore has no open item and no corresponding question, so calibration is scored on INFO, REASON, and BLOCKED decisions and excludes true-positive DONE decisions. H4, consistency at scale. On the fixed hundred-checkpoint, twenty-repeat set, Jev disagrees with itself on fewer than two percent of decisions, reported alongside the share of decisions that are consistently wrong. C′ is registered to sit between the typed-sample arm and Jev on this measure; the two-percent line applies to Jev alone. H5, counterfactual value. Replaying every legal action from saved checkpoints, C′'s chosen action carries lower mean regret than arm A's, with the gap concentrated in the INFO/REASON confusion and away from DONE/BLOCKED decisions, tested with the same case-clustered bootstrap as H1. Decision rules. H1 is supported if C′'s class accuracy exceeds both A's and B′'s with a confidence interval excluding zero. H2 is supported if C′'s verified-completion rate is at least A's and its premature-claim rate is lower, both with a confidence interval excluding zero, at expert calls no higher than A's. H3 is judged against the expected-calibration-error thresholds above; H5 is supported if C′'s mean regret is lower than A's with a confidence interval excluding zero. Two negative results are registered in advance as informative. A B′ that matches C′ would place the value in the typed-question pattern itself, which can run on commodity hardware. An A that matches C′ would mean typed questions add nothing a rule table lacked, and the vendor's own design documentation says such a result should be taken seriously. Phases and cost. J0, harness and generator selftests plus a smoke run with a deterministic mock worker, ran on a laptop at zero cost. J1, arms A, B, B′, and C′ with the real worker (the dual-Spark DeepSeek V4.1 Flash endpoint, PREREG-J1 addendum 4) across the dev, eval, and challenge splits, plus the consistency and replay runs, ran on the fleet's own Sparks at zero marginal cost. J2, the hosted Jev arm, ran through OpenRouter, with a probe call confirming the request shape before the full run, at a projected cost near one dollar. J3 covers expert consultation through a Bedrock model across every arm, reference arm D, and a sixty-case blind semantic adjudication sample. It runs under a fixed credits ceiling of sixty dollars, and an explicit rule limits the credits to judging and evaluation. They never pay for generating training content. The same constraint governed how this project's synthetic data was built, since case generation is template-based and deterministic from a seed, with no model involved. Adversarial review. The registration cleared only after three independent adversarial reviews, and the first two returned must not run. Along the way the reviews found and fixed a budget-accounting bug and a decision path that let CONSULT_EXPERT collapse onto the same underlying probability as CONTINUE . They also found splits where two of the four resolution classes were nearly impossible to get wrong. A mock worker had made two harder resolution classes unloseable by marking a document as satisfying a requirement whenever it was present at all, without checking entity or currency. Each was corrected and verified, with the splits rebuilt to eight hundred eval cases and two hundred challenge cases and every blocker tagged with the failure mode it exercises. The third review reproduced every prior fix independently, ran a Monte Carlo simulation against the actual eval-case distribution reaching roughly ninety-nine point eight percent power at the registered effect sizes, as PREREG-J1 addendum 3 records, and returned may run, with conditions: a cap of six thousand calls and five dollars for the decision head's own arm, a registered two-hundred-case Bedrock subsample inside the J3 ceiling, and one shared rule module, used by both the case generator and the mock worker, so that document currency and entity matching are judged identically everywhere they are asked. Arm C′, the self-hosted decision head, was registered in this same review, on the day TypeSafe stopped issuing new keys. What ran after J1. J1 answered H1 through H5 on the arms described above. Two follow-on registrations are reported alongside it in Section 6. J5 retrained the decision head on the same task, calling the result arm C″, added a separately fine-tuned "learned" head, arm L, and tested both on a fresh 800-case eval split (eval2) and the existing 200-case challenge split, under the same adversarial-review discipline as J1: a design review before any run, and an independent rescoring from the raw case files afterward. J6 and J7 registered a claim guard meant to catch the wrong-claim problem the challenge split exposed. J6's design failed its own pre-registered gate, and J7's first result was withdrawn after review found it circular. Section 6.5 reports why the guard was retired in favour of a fix to the pack builder. 6. Results J1 and J5 have run and were independently rescored from the raw case files by a party who did not build the scorer; every cell checked in that rescoring reproduced the launcher's own numbers within tolerance. The tables below report what was measured. J3, the expert-consultation and blind-adjudication phase, has not run and is reported below as an open item. The claim guard registered in J6 and J7 was retired (Section 6.5). 6.1 J1: the first registered comparison Table 1. Decision-class accuracy, eval split (n = 800 cases). measure A (rules) B (free LLM) B′ (typed, 8-sample, n = 400) C′ (zero-shot decision head) C (Jev, hosted) decision-class accuracy 0.575 0.527 0.583 0.545 0.540 C′ − A: −0.029, 95 percent CI [−0.060, +0.001], interval includes zero. C − A: −0.034, 95 percent CI excludes zero. Jev also ran on dev (0.560) and on the 200-case challenge split (0.536); C − A on challenge is −0.042, 95 percent CI excludes zero. H1 is rejected. Neither the zero-shot decision head nor the hosted service it was built to match beat the rule table on this split, and in the two comparisons where the interval excludes zero (C versus A, both splits), it lies on the side opposite the hypothesis. H2 asked whether acting on C′'s choice yields more verified completions and fewer premature claims per dollar than the rule table, at equal or fewer expert calls. H2 is rejected. C′'s verified completion falls below A's on eval by −0.126, with a confidence interval excluding zero, and on challenge by −0.060, again with an interval excluding zero. C′'s premature-claim rate on eval is higher than A's by +0.004, with an interval of 0 to +0.009. C′ uses roughly a hundred times A's expert calls on eval, 0.376 against 0.004 per case, and ninety-three times on challenge, against a registered prediction of equal or fewer. results/verification-J1-2026-09-24.md , lines 181 to 188, verifies these figures. H3 asked whether C′'s probabilities predict its own correctness, with an expected calibration error at or below 0.08. H3 is rejected for C′. Its expected calibration error is 0.5247 on eval and 0.5312 on challenge, more than six times the registered ceiling. B′'s sampled probabilities score an expected calibration error of 0.6286 on eval and 0.6265 on challenge, which confirms the half of H3 that predicted the local model would be uncalibrated. The Jev-specific reading of H3, on the hosted arm C, was never run. These figures are verified in results/verification-J1-2026-09-24.md , lines 190 to 197. 6.2 Consistency, replay regret, and the abstention band: H4, H5, H19, and H21 H4 asked whether Jev disagrees with itself on fewer than two percent of repeated decisions, replayed twenty times each on the same hundred fixed checkpoints per split. The two-percent line was written for Jev, the hosted reference, and Jev meets it. The open head was registered only to sit between the typed-sample arm and Jev, which it does. Jev disagrees with itself on 0.05 percent of repeats on both the eval and the challenge split. The open head disagrees on 4.95 percent of repeats on eval and 5.85 percent on challenge, inside the wider band an earlier addendum predicted for it and above the two-percent line. These J1 repeats ran at parallel two on an endpoint that a second campaign's accounts-payable run was sharing at the same time; the re-measurement below, under H19, shows what that sharing cost. Counted per decision point, Jev's twenty repeats are unanimous on 99 percent of points on both splits. The open head's twenty repeats are unanimous on 82 percent of eval points and 71 percent of challenge points, so on this reading Jev decides more consistently. The companion statistic registered alongside H4 measures whether those consistent answers are right. Of the points where all twenty of Jev's repeats agree, 65 percent on eval and 60 percent on challenge agree on the wrong action. For the open head the figures are 54 percent on eval and 51 percent on challenge. Where either controller answers consistently, the shared answer is wrong on more than half of those points. H5 asked whether Jev's chosen action carries lower replay regret than the rule table's, with a confidence interval excluding zero, where regret is the value of the best available action at a checkpoint minus the value of the action taken. Replaying every legal action from every saved checkpoint and scoring a case-clustered paired bootstrap shows Jev with the higher regret. Table 2. Mean replay regret by arm, case-clustered paired bootstrap, 10,000 resamples, seed 0. split A (rules) C′ (open head) C (Jev) C − A C′ − A C′ − C eval 0.013 0.115 0.069 +0.054 [+0.042, +0.066] +0.082 [+0.069, +0.096] +0.028 [+0.015, +0.042] challenge 0.111 0.196 0.177 +0.047 [+0.004, +0.093] +0.052 [+0.022, +0.082] +0.004 [−0.042, +0.049] Both intervals for C against A exclude zero, and both sit on the side opposite the registered prediction, with Jev's chosen action carrying significantly higher regret than the rule table's. H5 as registered is refuted. The instrument behind this comparison has a structural bias toward the rule table, separate from any scoring error. The deterministic rules controller completes every branch replayed from a checkpoint, including the branch the controller under test chose. For arm A, its own chosen continuation and the yardstick it is measured against are therefore the same policy from the second action onward. The share of checkpoints where the chosen action already equals the best-value action shows the effect, at 98.7 percent on eval and 88.9 percent on challenge for the rule table, against 93.1 and 82.5 percent for Jev, and 79.6 and 68.7 percent for the open head. Part of the rule table's low regret is therefore built into the yardstick. The refutation stands as arithmetic under the registered instrument, and J7 registers a replay in which the arm under test supplies the continuation policy in every branch, with no prediction carried over from H5. J7's H19 re-measured the open head's consistency under conditions J1 did not have: the same hundred checkpoints per split, replayed twenty times each, on an endpoint with no other campaign attached. Each split was run twice, once at one call in flight and once at four calls in flight (PREREG-J7 addenda 4, 5, and 7). At one call in flight the open head disagrees with itself on 0.05 percent of repeats on both the onboarding eval split and the new ap-eval split. That figure equals Jev's own J1 figure exactly. At four calls in flight the open head disagrees on 3.40 percent of repeats on onboarding and 1.95 percent on ap-eval. Table 3 reports all four cells, independently recomputed from the raw repeat files against the launcher's own numbers (PREREG-J7 addendum 8; results/verification-J7-H19-H21- 2026-09-29.md ). Table 3. J7 serving-consistency re-measurement, arm Cprime, 100 points times 20 repeats per cell, the same 100 points in both conditions within a split. split calls in flight mean agreement per-repeat disagreement onboarding (eval) 1 0.9995 0.05% onboarding (eval) 4 0.966 3.40% ap-eval 1 0.9995 0.05% ap-eval 4 0.981 1.95% Figure 3. How often a repeated decision changes. Share of repeats that disagree with the majority answer, twenty repeats per decision point. At four calls in flight the disagreeing repeats cluster on a single worker-thread residue class among the twenty draws. A thread pool batching calls together leaves that pattern, whereas a model answering differently each time would spread the disagreements evenly. Both concurrency-four figures sit below J1's per-repeat 4.95 percent on eval and 5.85 percent on challenge, measured under the shared endpoint. The inconsistency measured in J1 came from the serving stack. Ambient load from a second campaign sharing the endpoint caused it first, and batching added to it once concurrency rose. The shipped service now defaults to one call in flight. That default costs throughput, since the service serialises calls that could otherwise run in parallel, and leaves accuracy unchanged. H19's ap-eval condition required a fresh source run on arm Cprime that does not otherwise exist in this project. Scored on its own, that run reaches 0.533 decision-class accuracy on ap-eval against 0.648 for the rule table and 0.650 for the shipped head. It serves only as H19's source. The shipped head is served through the retrained C″ arm, so this run says nothing about it. H21 registered an abstention band meant to trade a small rise in handoffs for fewer wrong claims. The mechanically chosen band abstained on 1.8 percent of claims on ap-eval and 0.7 percent of claims on ap-challenge. Both figures sit below the run-to-run noise documented above for H19. In a live comparison against an independently rerun baseline, results moved by more than the band could have caused given how few claims it touched, and none of the cases whose claim status changed between the two runs were cases the band touched. H21 is reported as inconclusive. Its replacement is a counterfactual instrument that reads the band's effect from a single run's own trajectories, with no separately rerun baseline. 6.3 J5: the retrained head, fresh eval2 and the challenge split Table 4. Decision-class accuracy, J5 (n = 800 fresh eval2 cases; n = 200 challenge cases). split A (rules) C″ (retrained decision head) L (fine-tuned learned head) C (Jev, hosted) eval2 (fresh, n = 800) 0.566 0.728 0.701 not run on eval2 challenge (n = 200) 0.565 0.739 0.688 0.536 C″ − A, eval2: +0.161, 95 percent CI [0.131, 0.191]. C″ − A, challenge: +0.171, 95 percent CI [0.115, 0.224]. L − A, challenge: +0.124, 95 percent CI [0.033, 0.211]. C″ − C (Jev), challenge, paired on the same cases: +0.199, 95 percent CI [0.143, 0.254]. L − C″, eval2: −0.027, 95 percent CI [−0.056, +0.002]. Every interval above excludes zero except this last one. That comparison fails this project's own pre-registered acceptance rule for the learned head, so the learned head ships as experimental and the retrained head stays the default. 6.4 Wrong-claim rate and the scorer defect Table 5. Wrong-claim rate, challenge split (n = 200), registered exact-match scoring against a predicate-equivalent reading. arm registered (exact document id) predicate-equivalent A, rules 0.27 0.03 C, Jev 0.22 0.015 C′, zero-shot decision head 0.28 0.085 C″, retrained decision head 0.347 0.095 L, learned head 0.33 0.09 Error analysis traced 400 of the 435 wrong claims counted across every arm and split to one scorer defect, and an independent re-derivation from the raw case files confirmed it. The challenge generator attaches a second, independently valid document to a requirement that already has one on file, and the registered scorer's exact-id check marks the pack wrong for citing the other one. The predicate-equivalent column credits either valid document and scores every other mismatch, such as a stale document or a missing citation, wrong exactly as before. The defect is a measurement artifact in the harness and says nothing about how carefully the heads reason compared with the rule table. No decision-class accuracy number in Section 6.1 or 6.3 changes. Table 6. Wrong-claim rate, fresh eval2 split (n = 800), before and after a pack-builder fix. arm before, registered exact-id after, validated pack A, rules 0.0112 0.0088 C″, retrained decision head 0.0466 0.0113 L, learned head 0.0537 0.0163 Figure 4. Wrong-claim rate before and after the evidence-pack fix. Share of cases closed with a claim the registered scorer marks wrong. The 0.0466 reading sits above 0.04, the rate this project pre-registered as the trigger for redesigning the claim guard. A taxonomy of the judging head's 37 eval2 wrong claims, independently verified, found the cause. Thirty-one were packs the environment's own pack builder had padded with extra supporting documents beyond the ones the case key names; 91 of the 101 extra documents fail the harness's own validity predicate for entity and currency, the same check the ledger already applies elsewhere. Four were the worker citing the wrong document on an indirect requirement. The rule table makes that error on the same cases, so it is shared with the head. The remaining two were further citation mismatches. The taxonomy's independent verification traced none of the 37 to the head's own judgement question or to the elimination walk closing an item early. The remedy was registered before it was computed and independently verified afterward. The pack builder now checks every supporting document for entity and currency before listing it, using the predicate the ledger already applied to decide whether an item was met. Under the registered exact-id scoring rule this brings the eval2 wrong-claim rate to 0.0088 for the rules, 0.0113 for the judging head, and 0.0163 for the learned head, down from 0.0112, 0.0466, and 0.0537. Every case that changed moved from wrong to verified, and no other case moved. Both heads meet the pre-registered threshold of 0.02 on the validated pack and miss it on the original pack, and this paper reports both readings side by side. We found the fix by tracing why the threshold was first missed. It is reported as a correction to the environment, made and disclosed after the fact. 6.5 Retiring the claim guard Three approaches to a claim guard were tried after J5, and none earned a place in the design. A vote-based guard sampled several readings of the completion pack and voted. It separated wrong from verified claims at an AUC of 0.56, below its pre-registered gate. A deterministic evidence guard initially reported a 79-of-79 catch rate on the same kind of case. An adversarial review found that its entity and currency checks were the same two functions the case generator uses to define a wrong claim. A check scored against the rule it reuses cannot miss, so the number was withdrawn as circular. A reading-layer pilot then opened the cited document and compared it against the requirement, where the earlier guards had voted or checked dates. Combined with the deterministic layer, it produced numbers identical to the deterministic layer running alone, and on its own it separated a wrong claim from a right one at an AUC of 0.52. No model-based guard added anything past the deterministic check. All three guards are retired, and the deterministic validity check now runs inside Section 6.4's pack-builder fix, which validates every supporting document before a claim leaves. The check sits in the workflow engine's own bookkeeping and calls no model. Once it runs, the judging head claims wrongly at close to the rule table's own rate. 6.6 J6: transfer to a second process vocabulary J1 and J5 measured one synthetic process, supplier onboarding. J6 asked whether the rule table, the judging head, and the hosted Jev service carry to a second process built with the same generator family: accounts-payable exception routing, on two fresh splits, ap-eval at 400 cases and ap-challenge at 200 cases. The AP generator shares onboarding's blocker mix and document counts by construction. Before any AP data existed, adversarial review relabelled H15 and H16 as claims about vocabulary transfer, meaning a new process vocabulary layered on the same structure. Transfer to a process with a different blocker mix and obligation structure is registered as future work, J10. Table 7. J6 decision-class accuracy by split and arm, with the registered comparisons (case-clustered paired bootstrap, 10,000 resamples, seed 0). split A C″ L-onb L-ap C (Jev) C″−A L-onb−C″ L-ap−C″ L-ap−L-onb C−C″ ap-eval (400) 0.648 0.650 0.522 0.663 0.636 +0.002 [−0.029,+0.034] −0.128 [−0.174,−0.081] +0.013 [−0.016,+0.043] +0.141 [+0.088,+0.192] −0.015 [−0.052,+0.024] ap-challenge (200) 0.638 0.659 0.558 0.671 0.622 +0.022 [−0.023,+0.066] −0.101 [−0.158,−0.041] +0.011 [−0.034,+0.058] +0.113 [+0.038,+0.183] −0.037 [−0.088,+0.014] Figure 5. The head's gap to the rule table, run by run. Decision head minus rule table, decision-class accuracy. The first run used the zero-shot head. H15, rejected. C″ beat A by +0.002 on ap-eval, far short of the registered 0.03 bar, and the interval straddles zero; on ap-challenge the same comparison is +0.022, also short of the bar. The judging head landed inside its predicted band, 0.650 and 0.659 against 0.58 to 0.68. The rule table scored 0.648 and 0.638, above its own predicted band of 0.52 to 0.60. The predicted gap closed because A rose while C″ held, which leaves the rules level with the head on this vocabulary. J7 is registered to test whether the AP generator produces cases whose evidence the rule table can already see in full, leaving the judgement question nothing to add. H16, split. L-onb, the onboarding-trained head applied to AP with no retraining, came in at −0.128 against C″ on ap-eval and −0.101 on ap-challenge against a 0.05 band, with both intervals entirely on the wrong side of it. Its training on one process did not carry to the other. L-ap, retrained on 100 labelled ap-dev cases, recovered parity with C″ at +0.013 and +0.011, both inside the 0.03 band, and came in +0.141 and +0.113 above L-onb on the two splits, with both intervals clear of zero. On this evidence, a hundred of a company's own labelled cases (271 decision checkpoints here) brought the learned head level with the judging head on this vocabulary. H17, reference. Jev scored 0.636 on ap-eval and 0.622 on ap-challenge, at or below both the rule table and the judging head; only the ap-challenge gap against L-ap clears zero. Table 8 compares claim and handoff rates on packs rebuilt through the same validity filter on both sides. Table 8. J6 rebuilt-pack claim-rate and unnecessary-handoff comparison, Jev against the rules and the judging head, validated packs on both sides (case-clustered paired bootstrap, 10,000 resamples, seed 0). split comparison wrong-claim (predicate) diff [95% CI] unnecessary handoff diff [95% CI] ap-eval C − A −0.023 [−0.040, −0.005] +0.123 [+0.088, +0.158] ap-eval C − C″ −0.018 [−0.035, −0.003] +0.138 [+0.105, +0.173] ap-challenge C − A −0.055 [−0.095, −0.020] +0.125 [+0.080, +0.175] ap-challenge C − C″ −0.030 [−0.060, −0.005] +0.130 [+0.085, +0.180] Figure 6. Jev on the second process. Share of cases, Jev minus the judging head, with both sides scored on validated evidence packs. Left of zero means Jev does it less often. Jev claims a case done wrongly 1.8 to 5.5 points less often than the rules or the head, and hands a case off unnecessarily 12.3 to 13.8 points more often, the same trade seen on onboarding in Section 6.4. Its accuracy on this process sits at or below the other two arms, so the lower wrong-claim rate comes from caution. One caveat J6 raised is now closed. L-ap's training features and serving path both called C″'s own live judgement output, so its parity with C″ could have come from a classifier stacked on its own input. PREREG-J7 registered arm L0 to test this. It is the same learned head, trained on the same 100 ap-dev cases with every prior-answer feature removed, served with no call to the judging head. H22 ran on both AP splits and is independently verified in results/verification-J7-H22-2026-09-30.md . Table 7a. J7 H22, the learned head with no judging-head input, decision-class accuracy against the stacked head and the reference arms (case-clustered paired bootstrap, 10,000 resamples, seed 0). split L0 L0 − L-ap L0 − L-onb L0 − A L0 − C″ ap-eval (400) 0.660 −0.003 [−0.023, +0.016] +0.138 [+0.083, +0.189] +0.012 [−0.011, +0.034] +0.010 [−0.019, +0.040] ap-challenge (200) 0.670 −0.001 [−0.031, +0.031] +0.112 [+0.037, +0.181] +0.032 [+0.005, +0.061] +0.011 [−0.033, +0.055] L0 matches L-ap with no judging-head input, and the L0 − L-ap intervals straddle zero on both splits. On ap-eval L0 also claims wrongly less often than every reference arm, by 0.02 to 0.035 predicate points, with intervals clear of zero. A hundred labelled cases with structured and history features alone reach the accuracy of the stacked version, and the J5c and J6 circularity caveat on arm L is withdrawn. Read beside J8, the stacked classifier and L0 hold at 0.66 to 0.67 on the ap-dev split while an encoder fine-tune falls to 0.55. The encoder export carries the 201 checkpoints with open items from 58 of those cases; L-ap and L0 saw the same 201 plus the 70 terminal checkpoints that label DONE. On those same open decisions, feature design decided the comparison between the feature-based heads and the encoder. A second caveat still limits what J6 licenses. The four onboarding-harness arms, A, C″, L-onb, and L-ap, ran on the registered pack builder without the addendum-8 validity filter, while arm C ran through the isolated harness with that filter already applied. Table 8 rebuilds all four arms' packs through the same filter before comparing, and only that rebuild makes the claim-rate and handoff numbers readable. Decision-class accuracy stays as it was, since the rebuild replaces a claim's supporting-document list and leaves the chosen action alone. 6.7 J8: two open alternatives on the same benchmark Laya and CLM-8B, the two open decision models described in Section 2, were run on this project's own pre-registered benchmark, so that their measurement no longer rests on vendor claims alone. PREREG-J8 registered six hypotheses, H23 through H28, before any run. LY-ft is Laya fine-tuned on the training run the learned head used, exported as one row per open checkpoint: 1,567 six-way decisions and 2,901 per-item questions from 363 of the 600 onboarding training cases, and separately 201 decisions and 376 per-item questions from 58 of the 100 ap-dev cases. The cases the export leaves out were complete at their first checkpoint and carry no decision to learn from ( j8_export_training.py ; PREREG-J8 addendum 18). CLM-ft is CLM-8B with its two contrastive heads trained the same way. LY-0 and CLM-0 are both models as released, with no fine-tuning. Reference rows for A, C″, L, and Jev were already banked. Every arm answered the same rendered state the judging head answers. The case key, the generator's hidden blocker field, the oracle's expected action, and every judging-head prior answer were stripped before an encoder ever saw the text, checked by a selftest that greps the built training and serving text for those fields and asserts zero occurrences. Four splits ran: the two-hundred-case onboarding challenge split, the two-hundred-case ap-challenge split, the four-hundred-case ap-eval split, and a four-hundred-case seed-17 subsample of eval2. Each cell was independently recomputed from raw run-directory files within a tolerance of 0.0005, and each passed leakage, training-overlap and oracle-label audits before any hypothesis was read. Table 9. J8 decision-class accuracy, four splits, verified (addenda 13–16). split arm accuracy − A [95% CI] − C″ [95% CI] challenge (200) LY-ft 0.815 +0.251 [+0.191, +0.307] +0.076 [+0.042, +0.112] challenge (200) LY-ft seed 1 0.813 vs seed 0 +0.002 [−0.029, +0.043] challenge (200) LY-0 (untuned) 0.583 +0.018 [−0.002, +0.039] challenge (200) CLM-ft 0.687 +0.123 [+0.070, +0.176] −0.048 [−0.107, +0.010] challenge (200) CLM-ft seed 1 0.563 vs seed 0 −0.125 [−0.217, −0.029] challenge (200) CLM-0 (untuned) 0.537 −0.028 [−0.105, +0.051] ap-challenge (200) LY-ft 0.554 −0.084 [−0.156, −0.009] −0.105 [−0.173, −0.037] ap-challenge (200) LY-ft seed 1 0.560 vs seed 0 +0.006 [−0.062, +0.075] ap-challenge (200) LY-0 (untuned) 0.653 +0.015 [−0.016, +0.046] −0.006 [−0.053, +0.039] ap-challenge (200) CLM-ft 0.568 −0.070 [−0.140, +0.002] −0.091 [−0.151, −0.031] ap-challenge (200) CLM-ft seed 1 0.570 stable ap-challenge (200) CLM-0 (untuned) 0.656 +0.018 [−0.012, +0.047] ap-eval (400) LY-ft 0.539 −0.110 −0.112 [−0.166, −0.057] ap-eval (400) CLM-ft 0.586 −0.062 [−0.107, −0.016] −0.064 eval2 subsample (400) LY-ft 0.824 +0.276 [+0.232, +0.319] +0.069 [+0.049, +0.090] eval2 subsample (400) CLM-ft 0.641 +0.093 [+0.051, +0.137] −0.111 [−0.158, −0.065] Figure 7. Accuracy against the rule table on the harder set of each process. Each dot is one approach's difference in decision-class accuracy from the rule table on the same cases. The right-hand column gives the accuracy itself. Fine-tuning and retraining used the 600-case onboarding training split and the 100-case accounts-payable development split. References. Challenge: A 0.565, C″ 0.739, L 0.688, Jev 0.536. Ap-challenge: A 0.638, C″ 0.659, L-onb 0.558, L-ap 0.671, Jev 0.622. Ap-eval: A 0.648, C″ 0.650, L-onb 0.522, L-ap 0.663, Jev 0.636. Eval2 subsample: A 0.549, C″ 0.756 on 394 of 400 cases, L 0.735. H23 predicted LY-ft within 0.05 of C″ on eval2 and below the rule table on challenge. The prediction fails on three of the four splits, in the encoder's favour on two of them. LY-ft beats C″ by 0.076 on challenge and by 0.069 on eval2. On ap-eval it trails C″ by 0.11, far from the predicted near parity. The prediction's direction holds only on ap-challenge, the split whose fine-tune saw the least data, where LY-ft falls below the rule table and inside its own band. H24 predicted both untuned encoders near the majority-class baseline, and both land level with the rule table on the two splits they ran. H25 predicted CLM-ft below the rule table on every split. It beats the rule table on challenge and eval2, points the predicted way without significance on ap-challenge, and is supported only on ap-eval. H26 predicted both fine-tunes would claim a case complete wrongly more often than C″, and wherever the comparison is significant the reverse holds. Both fine-tunes claim wrongly less often than C″ on eval2 and ap-eval. On challenge LY-ft ties C″ on wrong claims, while CLM-ft shows a significant reduction. Neither fine-tune differs from C″ on ap-challenge. H27, descriptive only, holds. Laya answers in 0.17 seconds. CLM-8B averages 0.4 to 0.6 seconds on the onboarding challenge split and 0.6 to 0.8 seconds on the AP splits, with a tail reaching about nine seconds on both. The judging head's fast mode takes a median of 3.1 seconds per decision. H28, calibration, is confirmed, as reported below. Every reversal comes from the same class. On both splits where LY-ft beats C″, the whole advantage sits in REASON, with recall of 0.917 against C″'s 0.740 and the rule table's 0.104 on challenge, and 0.966 against 0.809 and 0.060 on eval2. With 1,567 labelled decisions, a 421-million-parameter encoder learns to recognise when a case needs judgement more accurately than the question-asking head infers it from the same rendered state. On the AP splits the same recipe, trained on only 201 labelled decisions, over-promotes INFO decisions to REASON at a rate of 0.53 for Laya and 0.39 for CLM-8B, against 0.02 and 0.06 untuned. INFO recall collapses and accuracy falls below the rule table it was meant to beat. Across the encoder's own two training sets, the 1,567-decision fine-tune beat the 201-decision one, so training size decided that comparison. The untuned encoders sit level with the rules on every split regardless of training size. Figure 8. Where the gains and losses sit, by decision class. Share of checkpoints of each true class that an approach classified correctly. The DONE class is left out because the harness decides it without a model call and every approach scores 1.000. Two calibration readouts close this section. H20, PREREG-J7 addendum 9, found the judging head's own item-level probabilities already calibrated on eval2, with expected calibration error of 0.0178 before any temperature refit and 0.0143 after one fitted on the training split. The 0.52 figure that started this line of work was a decision-class calibration reading, and that decision-class number stays open. H28 asked whether LY-ft's item-level calibration, after Laya's own vendor refit, beats C″'s uncalibrated reading on the same cases and comes within 0.05 of C″'s own refit reading. On the same four hundred eval2 cases, LY-ft's item-level ECE is 0.0057 uncalibrated and 0.0054 after its refit, against C″'s 0.0313 uncalibrated on the identical cases. Both registered conditions hold, and H28 is confirmed. The head was already well calibrated at the item level before this campaign ran, so Laya's margin over it is small in absolute terms. Figure 9. What training on a process's own decisions did. Decision-class accuracy on the harder set of each process, 200 cases each. Table 10. J8 scorecard, all splits verified (PREREG-J8 addendum 17). hypothesis onboarding challenge eval2 (400) ap-challenge ap-eval H23 LY-ft vs C″ wrong, encoder +0.076 wrong, encoder +0.069 supported (below A) wrong, trails by 0.11 H24 LY-0 near random wrong (level with A) not run wrong (level with A) not run H25 CLM-ft below A wrong (+0.123) wrong (+0.093) direction right, not significant supported H26 wrong claims above C″ wrong (tie) reversed (fewer) wrong (no difference) reversed (fewer) H27 latency Laya 0.17 s, CLM 0.4–0.6 s onboarding / 0.6–0.8 s AP, heavy-tailed to 9 s, reported H28 calibration confirmed Four limitations, registered in PREREG-J8 addendum 17, qualify this section; Section 7 reports them together with every other campaign's. 6.8 What was not measured J3, the expert-consultation phase and its blind-adjudication sample, has not run. This project's registered calibration function cannot score the learned head, because the head outputs a classifier's class probabilities and the function reads a per-item probability map. A design-time reading exists, but it was fitted on the same data used to build the model, so this paper leaves it out. Both gaps are reported as not yet measured, which carries no negative result. 7. Limitations Every case in this study is synthetic, generated deterministically from templates with no model involved. That gives ground truth correct by construction and guarantees that no client evidence appears anywhere in this project. The cost is that we do not yet know whether the same distinction holds on the messier documents a real supplier file contains. Eval2 came from the same case templates and process family as the split the retrained head was tuned against. Its gains (Section 6.3) are a fresh sample from the same regime, and nothing here says whether they survive a different workflow. J6's transfer test carries the same limit to a second process. The AP splits share onboarding's blocker mix and document counts by construction, so the transfer results in Section 6.6 concern a new process vocabulary layered on the same structure. Whether these results hold on a process with a different blocker mix and obligation structure is registered as future work, J10. The retrained head's wrong-claim rate cleared its pre-registered bar only after a defect in the environment's own pack builder was found and fixed. On the original pack it remains above the trigger. Because the fix came after the trigger was missed, both readings are reported side by side (Section 6.4). The learned head ships as experimental by the same pre-registered rule, having failed its own acceptance test against the retrained head it was meant to match. No model-based claim guard passed a fair test. The vote guard and the reading-layer pilot each added nothing beyond a deterministic check that reproduces the benchmark's own validity rule, the same check the pack-builder fix relies on (Section 6.5). The worker and the decision head in arm C′ are both built on DeepSeek V4.1 Flash, served through a three-bit-per-weight quantization spanning two GPUs as one model. Another quantization, another base model, or an unquantized reference may calibrate and discriminate differently, and this registration says nothing about where on that curve the effect appears or disappears. The one-off probe and the service's own validation smoke, both against qwen3.8:27b (Section 3), showed only that a logprobs read is possible on a model of that kind, and J1 does not measure that model. The decision head reads a single generated token per question, a narrow window onto what a model believes. A question whose answer needs more than one token cannot be asked this way, and this design makes no attempt to widen the window. The calibration measurement has a gap that was registered in advance. A correctly timed completion closes every obligation and produces no corresponding question, so H3 measures calibration on open and blocked cases and leaves out the moment a case correctly closes. Adversarial review also made a related structural fact explicit. In this version of the question set, CONTINUE and CONSULT_EXPERT are decided from one shared probability, so a controller cannot yet state its confidence that a case needs a stronger model separately from its confidence that it can handle the case itself. A question set that separates them is future work. The counterfactual replay behind Section 6.2's H5 result carries a structural bias. replay() completes every branch from a checkpoint, including the one chosen, with the deterministic rules controller regardless of which arm is under test. Arm A benefits from this every time, because its continuation and the yardstick are the same policy. The H5 refutation holds as arithmetic under this instrument and ranks decision quality poorly. J7 registers a replay that completes each branch with the policy of the arm under test. Section 6.2's consistency figure depends on the serving configuration as well as on the model. J7's H19 traced the open head's apparent inconsistency against Jev to the serving stack, first to ambient load from a second campaign sharing the endpoint and then to batching once concurrency rose. At one call in flight, on an endpoint with no other job attached, the open head disagrees with itself on 0.05 percent of repeats, matching Jev's own J1 figure exactly. The shipped service now defaults to one call in flight, which removes the gap at the cost of throughput; this project did not measure how far concurrency can rise before the figure moves, or whether Deliberate mode, Section 4's higher-cost setting that samples several answers per decision instead of one, closes the gap faster than serialising calls does. H21's abstention band, registered as a partial hedge against the disagreement that remains, abstained too rarely to be read from a live rerun against a separate baseline; whether it helps at all is still open, and a counterfactual reading of a single run's own trajectories is registered in its place. A typed-question head should be kept away from arithmetic and date logic. The vendor's own documented failure modes for Jev name both, and this experiment's environment applies the same lesson by keeping currency and entity checks in a shared rule module that no model, hosted or local, is ever asked to reproduce. J8's run against the two open encoders carries limits of its own, registered in PREREG-J8 addendum 17. Every case in J8 comes from the one synthetic generator family this project has used throughout, and the training labels and the scoring function share one oracle, the same limit the learned head carried in Section 6.6. The DONE class is a zero-call shortcut credited to every arm alike, encoder and rule table. The encoder recipes were reconstructed from the vendors' own published primitives without running vendor tooling end to end, and CLM-8B's encoder runs behind a hand-written transformers shim whose numerical parity with the vendor's own vLLM serving path is unverified. The runs hit three Metal-backend crashes on the serving side before a clean run. That was a serving-stack problem with no bearing on the modelling, of the same kind J7's H19 traced on the fleet's own endpoint. A later diagnostic, registered in PREREG-J9 and reproduced independently, bounds what the encoder learned. A four-feature rule computed from the rendered observation alone, with no learning and no document reading, drives the same elimination walk to 0.754 on the eval2 200-case subsample and 0.793 on the challenge split, between 0.02 and 0.045 below the fine-tuned encoder depending on the split. Part of the encoder's gain is a cue the state already carries, and J9 is registered to measure how large that part is. A second late diagnostic examines the shared oracle itself (PREREG-ORACLE, 6 October). A frontier model that had seen neither arm's answers judged 197 blind checkpoints from the case files alone. Where the oracle's rule and the file give one answer, it agreed every time, 28 of 28 INFO and 83 of 83 REASON. On the 120 checkpoints where the rule table and the head disagree, it sided with the oracle on 99, and the oracle's class there was the head's in 85 cases and the rule table's in 8. It never chose BLOCKED, because the fact that no counterparty can supply a document lives in the hidden key and is absent from the file. It also read every closed item that rested on a related-entity document as still needing interpretation, a difference in process definition with no model fault behind it. The same strict reading makes it agree with the registered scorer that 36 of 37 eval2 wrong claims were wrong. On the facts it agrees with Section 6.4's taxonomy, which found most of the extra documents in the padded packs failed the validity predicate; the two differ only on whether a related-entity document closes an item by interpretation, which the process rules allow and the adjudicator read strictly. Three checkpoints where the oracle itself may be wrong are flagged for a human look. In J8 training size decided the encoder-against-encoder comparison and feature design decided the head-against-encoder comparison, and a learned mode must be checked on held-out data before it is promoted. The judging head remains the default. It reads the cited evidence, which neither encoder does, and it ships with a Jev-shaped API and an audit row behind every decision. 8. Reproducibility Every module in this project (case generator, environment, controllers, scorer, runner, and the decision head itself) is written in Python 3.11 against the standard library only, so the harness runs anywhere a plain Python install does. Ground truth is correct by construction. The case generator builds the hidden key first and then renders the documents, emails, and application data from that key, so no data is labelled after the fact. The frozen question set and decision table carry a recorded SHA-256 hash checked at run time ( questions_v1.json , version 1.2), and the harness refuses any run against a different question set. Every adversarial review is pinned to a specific commit, and every module carries its own self-test, printed as an N-of-N pass count, that must pass before its output is trusted downstream. The worker model is pinned per run in runs//config.json ; a change of worker, such as the move to DeepSeek V4.1 Flash, starts a new run so that no run mixes two workers. Every phase reported in Section 6 was independently verified from raw run files by a party who did not build the scorer, in each case importing the scorer's own metric functions without re-implementing their logic. J6's transfer arms, Jev's read-out, and the rebuilt-pack rescoring behind Table 8 are each verified in their own file ( results/verification-J6-2026-09-26.md , results/verification-J6-armC-2026-09-27.md , results/j6-rebuilt-packs-2026-09-27.md ). Section 6.2's H4 and H5 numbers were verified twice, the second time by direct recomputation from the raw repeat and branch files ( results/verification-J1-H4H5-raw-2026-09-28.md ); every cell matches the numbers reported in Section 6.2 within the stated tolerance. H19 and H21 were verified the same way, with all 28 of 28 cells matching the launcher's own numbers within tolerance 0.0005 ( results/verification-J7-H19-H21-2026-09-29.md ). H22 was verified by recomputing every cell from raw trajectory files and checking the model's own provenance fields, hash, feature list, and fast-head call count ( results/verification-J7-H22-2026-09-30.md ). J8's four splits are each independently verified in their own file ( results/verification-J8-challenge-2026-09-29.md , results/verification-J8-ap-challenge-2026-09-29.md , results/verification-J8-ap-eval-2026-09-29.md , results/verification-J8-eval2-2026-09-30.md ), each within tolerance 0.0005 and each additionally audited for leakage, training-set overlap, and the oracle-label limitation before a hypothesis was read. Section 6.7's two calibration readouts are each in their own file ( results/j7-h20-calibration-2026-09-30.md , results/j8-h28-calibration-2026-09-30.md ). Ethics and data No client data of any kind appears anywhere in this project. Every case, document, and email is generated locally and deterministically from a synthetic policy and synthetic supplier templates. A hosted model is called in only two places, a narrowly scoped expert-consultation step and a blind adjudication sample. Both are funded under a fixed credits ceiling, and an explicit rule restricts them to judging and evaluation, so neither generates any of the material the system is tested on. Those two calls are the only traffic that leaves the local network, apart from calls to the hosted Jev service through OpenRouter, which are logged and reported as a separate cost line. Acknowledgments The author's automated research fleet built the harness, ran the three adversarial reviews, and implemented and self-tested the decision head, working from a brief and two internal working papers. The internal papers are cited here only as the origin of the question and are not public sources. ### Building a Jev Lookalike in One Evening URL: https://dinand.com/research/decision-head/blog/ Section: Research · Decision Head · Blog post (23 September 2026) Building a Jev Lookalike in One Evening I had been designing an experiment around a product called Jev, from a company called TypeSafe, and when I went to open an account, TypeSafe was not issuing new API keys at the time. I reached Jev through OpenRouter for the comparison instead. Jev does something I think gets too little attention when people talk about AI in business. You give it the state of a case and a list of typed questions, yes-or-no or multiple-choice, and for each question it hands back a calibrated probability with no prose attached. Plenty of the decisions that sit inside a governed process suit that shape better than a chat response does. Does this case have the evidence it needs, does it need a document nobody has yet, should someone reason harder about what is already on file, is this case done? Each of those wants a probability and an audit trail, and a paragraph of prose gets in the way of both. What building it myself buys, beyond spite Being locked out pushed me toward a question worth asking whether or not TypeSafe ever reopens enrollment, which is what a company gets by running this in-house. I count nine things. Everything stays inside the network, so a pilot can start before a data-processing agreement is signed. The confidence numbers are calibrated on your own labelled cases, and you can hand the reliability table to your own auditor. Pointing the same service at a new process takes a question file and a calibration run, with no new contract to sign. Once you have banked enough of your own decisions, you can distil a smaller model trained on exactly those cases and measure what that buys you. Every answer carries an audit row holding the exact question, the raw probability, the model and the latency, so a decision from six months ago can be reconstructed in full. The same service runs unchanged on anything from a small local model up to a frontier model in your own cloud account. The whole premise is that each question gets answered in isolation, and the service tests that claim on your own data, so the evidence for it comes from your own cases. The running cost is whatever your own hardware already costs. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head , after two independent security reviews, so your security team gets to read every line before it touches a real case. A purpose-built model may still win at the job it was trained for, and running my own means I can measure the two side by side, which is what the rest of this post does. I had already been reading two internal analyses of where Jev fits inside a governed workflow, along with a study my colleague Thordur Arnason had published on LinkedIn. He measured Jev against a frontier model at reviewing agent trajectories after the fact, on a benchmark called tau-squared-bench, and he reported what he found straight. Jev came close to Haiku 4.5 at flagging wrong database changes, four times faster, at roughly four percent of the cost, and open-weight models running on one of our own DGX Sparks scored no better than a coin flip at the same task. He was also open about a blind spot in his own study. An agent can fail by doing nothing wrong and still doing too little, and a review of recorded actions will never catch that, because there is no action to point at. That gap, the question of what should happen next, was the one I wanted to test, and then the account page told me I couldn't. The probe Before writing the experiment off, I made one call to our own qwen3.8:27b model through Ollama's native chat route, with logprobs turned on, and asked it a single yes-or-no question. I wanted to see whether the model would hand back anything beyond a token. A probability came back in that one call. The model's belief about its answer was in the response all along, and most code that calls these models day to day ignores it. A model already computes this number whenever it generates a token, so the experiment no longer hung on one vendor. Jev's trick is to ask one narrow question at a time and read the answer back as a probability, and anyone can reproduce that. TypeSafe trained its own model specifically for this. The pattern around the model needs a shared state, a single question per call, one token read as a distribution, and a calibration step fitted on your own cases, and any team can build it. The service's own smoke test, run the same night, makes a better illustration than that single probe, because it asks all three question types about one realistic case at once. Treat it as an illustration only, since it was never meant as a benchmark. Below are the request and the response, with every number exactly as it came out of the run. POST /v1/systemone { "state": "Acme Logistics BV, a subsidiary of Acme Holding NV, submitted a certificate of insurance dated 2026-01-15 naming Acme Holding NV as the insured party rather than the subsidiary. No endorsement schedule or subsidiary rider was included.", "model": "jev-latest", "questions": { "insurance_satisfied": { "type": "noul", "instructions": "The evidence on file satisfies the insurance requirement for Acme Logistics BV." }, "next_step": { "type": "choice", "instructions": "What should happen next?", "options": ["CONTINUE", "REQUEST", "HANDOFF"] }, "urgency": { "type": "score", "instructions": "How urgent is this case?", "levels": ["low", "medium", "high"] } } } 200 OK { "model": "qwen3.8:27b", "answers": { "insurance_satisfied": {"noul": 0.001}, "next_step": {"choice": "REQUEST", "probabilities": {"CONTINUE": 0.007, "REQUEST": 0.910, "HANDOFF": 0.083}, "confidence": 0.827}, "urgency": {"score": "medium", "probabilities": {"low": 0.188, "medium": 0.412, "high": 0.400}, "confidence": 0.012} } } Every number above comes from the nine calls that made up the smoke test. The service spotted the entity mismatch and said the certificate does not satisfy the requirement, with 99.9 percent probability. It judged the mismatch fixable and asked for the right document, keeping the case out of a person's queue, choosing REQUEST with probability 0.91 and a margin of 0.83 over the next option. On urgency it landed on medium by a hair, with low, medium and high almost evenly split and the model about as unsure as a model can be. An even split like that still carries information, because it says the model doesn't know, and calibration is there to keep that signal intact. The first call took 9.4 seconds, and a repeat with the same state already warm in the prefix took 2.7. To check the isolation claim directly, I asked every question a second time with the other two visible, and none of the three answers changed. The satisfied and request readings are logged in the fleet's governance log for 23 September 2026, wiki/governance/log-2026-09-23.md . The model field in the request is an alias that lets a client keep its existing request exactly as it was, and the response always names the model that answered. Three reviews, two refusals Everything in our research process now goes through an adversarial review before it runs, and this design went through three rounds of it. The first round found nine problems. One was a scoring bug that quietly double-counted a budget. Another was a splits design in which two of the four outcome classes were nearly impossible to get wrong, which would have drained the test of any meaning without anyone noticing. The reviewer's verdict was that it must not run. We fixed each problem with its own test, rebuilt the evaluation set larger and harder, and sent the design back. The second review caught something worse. The cheap mock worker we use for early testing had been marking a document as satisfying a requirement whenever it was present, without checking whether it named the right company or was still valid. With that shortcut in place, a wrong document on a real case would have gone straight through. We fixed the worker and spot-checked the real model against cases built specifically to expose that failure, and the model got every one of them right. The second reviewer still blocked the run. The third review confirmed every fix independently and ran a statistical check against the evaluation data, which showed the design had enough statistical power to detect the effects it is looking for. It approved the run with conditions. We had to cap spending and add a shared rule module so two parts of the system can't quietly disagree about what "current" means. We also had to register the open decision head itself as a formal arm of the experiment, next to the rules-based approach, a free-choosing model, and the same local model asked Jev's exact questions. A separate piece of work landed the same night. DeepSeek V4.1 Flash went into production on a shared endpoint that runs across both of our DGX Sparks as one model, serving 27 tokens a second on a single stream and 65 aggregate across six sessions, per the throughput ladder banked in experiments/dual-spark-tp2/results/verification-candidate2-2026-09-23.md . It became the assistants' own primary model, and it took over as the worker behind this experiment too. qwen3.8:27b , which answered the probe and the smoke test above, was the model I had on hand that night to prove the mechanism worked. The experiment itself runs on DeepSeek now, recorded as a fourth addendum to the registration. The speed question The open head takes about three seconds where Jev takes a third of one, and these are the measurements behind both figures. Jev, through OpenRouter, answered a one-question probe in 0.35 seconds when I measured it on 24 September 2026. The call cost $0.0000128 for 304 input tokens. On the same day, the open head's fast mode ran DeepSeek V4.1 Flash on our two DGX Sparks while four benchmark streams were already hammering the same hardware. PREREG-J6.md and results/design-review-J6-2026-09-24.md record these figures: measure value median 3.1 s p90 4.8 s No worst-case latency or request count was banked for this run. Jev wins that comparison outright. Deliberate mode samples eight answers per call for the harder decisions, so it will cost roughly eight times as much per call, and I haven't measured its latency yet. A gate like this fires once per checkpoint in a multi-step process. In the J5 eval2 run, 794 cases passed through 2,681 checkpoints, about 3.4 per case. The steps on either side of a checkpoint, somebody chasing a supplier for a document, say, or routing a case to an expert, take anywhere from minutes to days. Against a two-day wait for a document, three seconds won't register with anyone. Seconds, logarithmic scale. On 800 eval cases, the zero-shot open head scored 0.545 decision-class accuracy, against 0.575 for the rule table it's supposed to beat. The gap is -0.029, and the interval around it runs from -0.060 to +0.001, so I could not rule out that the rule table was better. Jev's own run landed the same way, at 0.540 on that eval set, 0.560 on dev and 0.536 on the 200-case set built with contradictory evidence, below the rule table on both, and this time with an interval that excludes zero. On the contradictory-evidence set, every arm, rules included, wrongly called a case complete 21 to 28 percent of the time. The test's own scoring caused most of those errors. The test sometimes plants a second, equally valid document and then marks a citation of it wrong. Once that was fixed, the rate dropped to between 1.5 and 9.5 percent depending on the arm. Then I retrained the head and ran it again on 800 fresh cases. It scored 0.728 against 0.566 for the rule table, and 0.739 against 0.565 on the harder set, where it beat Jev's own 0.536 on the identical cases by 0.199, measured on the 199 cases both scored. Every one of those gaps excludes zero. It was the first result in the project where the open head beat both the rule table and Jev. On the same fresh set, though, the retrained head wrongly called a case complete at a rate of 0.047, which is above the line I set in advance as my trigger to redesign the safeguard meant to catch that mistake. I've built and tested three guards since, two of them model-based. The first looked perfect, catching 79 of 79, until a review found it had been scored against the same rule it was supposed to be checking, so it could never have missed. I threw that number out. A second, a vote across several readings of the case, separated wrong claims from right ones at an AUC of 0.56. A third read the cited document itself and did no better than a plain deterministic check. All three are retired, and the fix that worked went into the pack builder, which I cover under "Where this stands". In return for the slower answer, the state of a live case stays inside our own network, and every answer above came with an audit row I can hand to someone six months from now. Retraining the head on a hundred of our own labelled cases doesn't need a GPU at all. Jev is faster and needs no infrastructure of its own to run, and it was tuned on more data than we'll log ourselves any time soon. The retrained head now beats the rule table on two test sets. Against Jev it won on the harder set, the only one where both ran in that round. Its wrong-claim rate was the problem still open after those runs, and the pack-builder fix later in this post is what dealt with it. How consistent any of this is I finally scored two questions I'd left blank since day one. One asks how often each approach disagrees with itself. The other asks whether its chosen move beats the alternatives once you replay the case forward. On the first pass Jev disagreed with itself on 0.05 percent of repeats and the head on 4.95 to 5.85 percent, which looked bad for the head. A second job of mine had been hitting the same endpoint during that pass, though, so I didn't trust the number. I re-ran the head alone on the same points with nothing else on the endpoint, first one call at a time and then four. With one call at a time, the head disagreed with itself on 0.05 percent of repeats, which matches Jev's own figure exactly. At four calls at a time, disagreement climbed to 3.4 percent, and the disagreements piled up on one worker thread, over and over. That clustering is the fingerprint of calls being batched together on the server. The inconsistency came from my serving setup, so the service now defaults to one call at a time, and the only price of that default is speed. The repeats also showed that where every repeat agrees, both approaches agree on the wrong answer more often than on the right one. Share of repeats that disagree with the majority answer, twenty repeats per decision point. I also tried having the service abstain on the claims it is least sure of, in the hope that abstention alone would bring the wrong-claim rate down. It abstained so rarely that an ordinary rerun moved the numbers more than the abstaining did. I'm calling that test inconclusive. I have registered a cleaner version that reads the effect from a single run's own trajectory instead of comparing two separate runs against each other. The replay question came back against Jev. The replay, though, finished every branch, the rule table's own included, by running the rule table forward, which stacks the deck before either side makes a move. I'm rebuilding that test and won't report it as a fair ranking. Where this stands The service is built and has now run twice, and both runs were rescored independently from the raw case files by someone who didn't build the scorer. It is a few hundred lines of Python that use nothing beyond the standard library. It runs against the DeepSeek endpoint on our own hardware, and the same code can point at a small model on a workstation or a frontier model in a cloud account. It calibrates on labelled cases you supply and tests its own claim that questions stay isolated from one another, and every answer gets an audit row. The prediction I was least attached to came true. On the first evaluation set the open model asked Jev's exact questions scored 0.545 and the purpose-built Jev 0.540, and neither of them beat a plain rule table. Retraining the head on the project's own cases changed the answer. On a fresh eight hundred cases it beat the rule table by 0.161, and on the harder set it beat Jev by 0.199. I also ran a smaller, separately fine-tuned version of the head, the distillation step I said back then was still ahead of me. It scored 0.701 to the retrained head's 0.728, too far apart to call the two equivalent under the rule I set in advance, so it ships labelled experimental and the retrained head stays in place. The retrained head's biggest number came with a cost. It wrongly called a case complete more often than the line I set for myself before I saw a single result. I built and tested three guards for that problem, two of them model-based, and none of them worked. I threw out the first because a review found it was scored against the same rule it was supposed to be checking. The second, a vote across several readings, reached an AUC of 0.56. The third opened the cited document and read it against the requirement, going further than any vote or date check, and it failed in the same way. Combined with a plain deterministic check, its numbers came out identical to that check running alone, and by itself it separated a wrong claim from a right one at an AUC of 0.52. No model-based guard added anything past the deterministic check. The pack builder now runs that same validity check, with a caveat. On my own test cases it catches wrong claims by the same rule my own generator used to label them wrong in the first place, so a real deployment would need a rule of its own. Then I went looking for the reason the wrong-claim rate was high at all, so I could explain it as well as report it. Of the retrained head's 37 wrong claims on the fresh set, 31 were packs where the environment listed extra supporting documents it never checked, most of them decoys built to catch a careless citation and never meant to pass. Four were the worker citing the wrong document on a hard case, a mistake the plain rule table also makes on the identical cases, and the last two were further citation mismatches. An independent check traced none of the 37 to the head's judgement question or to its elimination walk. I fixed the pack builder to check every document the way the rest of the system already checks them, and re-ran the numbers: rules: from 0.0112 to 0.0088 retrained head: from 0.0466 to 0.0113 learned head: from 0.0537 to 0.0163 Every case that changed moved from wrong to verified, and nothing else moved. Someone else checked all of it independently before I trusted it. The fixed reading meets the line I set for myself in advance and the original reading misses it, and I'm reporting both so the fixed one doesn't quietly replace the other. I also found the defect only after the line had first been missed. Share of cases closed with a claim the registered scorer marks wrong. Onto a second process Every number above came from one synthetic process, supplier onboarding, so the obvious next step was to see whether any of it carries over to a different one. I ran the rule table, the judgement head and Jev on a second synthetic process built the same way, accounts-payable exception routing. On the new process the rule table scored 0.648 on ap-eval and the judgement head 0.650, a gap well inside noise. Decision head minus rule table, decision-class accuracy. The first run used the zero-shot head. Pointed at the new process with no retraining, the learned head dropped to 0.522 on ap-eval. Retrained on a hundred labelled cases from the new process itself, 271 decisions in all, it caught back up to the judgement head. Jev scored below both the rules and the head on the new process. When I looked at the kind of mistakes it made, Jev claimed a case done incorrectly less often than either of the others and handed cases off unnecessarily more often, a cautious habit that cost it accuracy here. Share of cases, Jev minus the judging head, with both sides scored on validated evidence packs. Left of zero means Jev does it less often. A caveat in the first version of this section has since been settled. When I first wrote this up, the learned head's training features and its serving path both called the judgement head's own live output, so its parity with the judgement head could have been a classifier stacked on its own input. I registered a follow-up to check that directly. I stripped out every judgement-head feature, trained on the same hundred labelled cases, and served the result with no call to the judgement head at all, and it still matched. It scored 0.660 on ap-eval to the stacked version's 0.663, and 0.670 on ap-challenge to 0.671, both gaps well inside noise. On ap-eval it also claimed wrongly less often than every other arm. A hundred labelled cases and the structured features were enough by themselves, which retires the stacking caveat. The encoder I fine-tuned on the same split fell to 0.55, while the stacked head and this independent one both held at 0.66 to 0.67. The encoder's export keeps only the 201 decisions where something was still open, from 58 of those cases. The learned head saw those same 201 plus 70 closing checkpoints that label a case done. On those same open decisions the feature-based heads still beat the encoder, so in that comparison the feature design decided the outcome. All of this rests on one new process built the same way as the first, and I still don't know whether it holds on a process built along different lines. Two open models showed up, and I ran them anyway Two free open encoders, Laya and CLM-8B, appeared within days of everything above. Both are both small, and neither reads evidence. I registered predictions before running either, as I do for everything else in this project, and I was wrong more often than I was right. Fine-tuned on 1,567 labelled onboarding decisions, the open ones from the same training run my learned head used, Laya scored 0.815 on the harder set of the process it learned, against 0.739 for my judging head and 0.565 for the rule table. I had predicted near parity with the head and a loss to the rules, so both predictions missed in the same direction on the same split. Each dot is one approach's difference in decision-class accuracy from the rule table on the same cases. The right-hand column gives the accuracy itself. Fine-tuning and retraining used the 600-case onboarding training split and the 100-case accounts-payable development split. Then I ran the same recipe on 201 labelled decisions from the accounts-payable process and expected the same story. Laya fell below the rule table it was supposed to beat. Across the encoder's own two training sets, with the model and the recipe held fixed, the result flipped, and in that comparison I put the flip down to the amount of labelled data. Two hundred decisions left it with a habit of promoting requests to consultations it didn't need to make, which it had not picked up from some 1,600. Share of checkpoints of each true class that an approach classified correctly. The DONE class is left out because the harness decides it without a model call and every approach scores 1.000. Untuned, both models sat level with the rules and moved nothing in either direction. I also checked calibration, since Laya's maker leans hard on that number. The first version of the head, tested back in J1, was badly calibrated at the decision level, but at the item level my own head turned out to be well calibrated already, before I did anything about it. Both item-level errors were small, the open model's a little smaller than mine, and that edge was less than I'd expected going in. The head keeps the job because it reads the document and writes a ledger row I can hand to someone later, which the open encoders can't do. ### The Measured Enterprise 2026 URL: https://dinand.com/research/measured-enterprise-2026/ Section: Research · Annual report (first full draft, 6 October 2026) Tinholt Lab Annual report · 2026 Pre-registered measurement of delegation and oversight The Measured Enterprise 2026 Organisations are handing decisions to AI agents faster than they are building the oversight to match. The lab measures where delegation breaks and what oversight buys back, in pre-registered campaigns on local models, and publishes the result whichever way it falls. 90,880 organisation designs swept for the price of delegation depth 30,000 supervised episodes across five irreversibility tiers 18,360 units audited across all fault classes, including 3,240 real diaries 100 simulation runs of first-mover AI adoption across four industry verticals 0 mismatches when 14,950 convex and 14,952 linear scored episodes, 29,902 in all, were recomputed independently Dinand Tinholt Head of AI Center of Excellence, Americas, Capgemini Draft First full draft · 6 October 2026 The Measured Enterprise 2026 Contents Summary Executive summary 3 Chapter 1 The Delegation Index 5 Chapter 2 How the lab works 9 Chapter 3 AGENESIS: first-mover adoption 11 Chapter 4 Error Independence 12 Chapter 5 Delegation Cliff 14 Chapter 6 SIGIL 16 Chapter 7 AGENESIS-2: sprawl cost 18 Chapter 8 Trial Balance 20 Chapter 9 The open decision head 22 Chapter 10 What comes next 26 Chapter 11 Corrections logged 27 Appendix A Measured claims index 29 Appendix B Glossary 32 Appendix C Sources 34 Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Charts are rebuilt from those numbers. 06 OCT 2026 Summary Executive summary The reading The Delegation Index reads 47.2 for the third quarter of 2026, with a band of 40.8 to 53.2. Its inputs are the lab's own synthetic measurements, the simulations and controlled model runs on generated tasks from five published campaigns, 139,646 units in all. It says nothing about the market or about any single firm. The number summarises how those campaigns came out when enterprise work went to AI agents under oversight and the oversight budget stayed fixed. At a reading of 50 the two balance under the conditions the lab measured. 47.2 Delegation Index, Q3 2026. Band 40.8 to 53.2. Five campaigns, 139,646 units. This first reading sits 2.8 points under balance, and its band contains 50. This quarter's evidence cannot tell balance from imbalance, which is why the number is always quoted with its band. Five components Each component reads one published number and scores it from 0 to 100. All five weigh the same. Component Campaign Published number it reads Score Band Delegation price Delegation Cliff depth −1.499 against reviewer capacity +1.456 49.3 none published Oversight yield SIGIL value of plain review against no oversight, five tiers 46.9 46.8 to 47.0 Audit detection Trial Balance checklist reader at a full read, 0.583 58.3 47.9 to 66.7 Reviewer independence Error Independence mean cross-family error correlation, −0.102 55.1 36.0 to 73.5 Sprawl AGENESIS-2 cost exponent under a flat monitoring budget, 1.4665 26.7 23.9 to 29.5 Sprawl holds the index down. With the monitoring budget flat, oversight cost grows as fleet size to the power 1.47, and that leg scores 26.7. Audit detection lifts it, because a reviewer who read the agent's whole record with a per-step checklist found 0.58 of the skipped steps. The other three land within five points of balance, and Chapter 1 walks through the arithmetic step by step. What the campaigns found Error Independence ran one model on two different machines and found that the copies share their mistakes, with an error correlation of +0.447 on the items where the models disagreed. Pairs from different model families averaged −0.102. Delegation Cliff swept 90,880 simulated organisations. Each added level of delegation depth cost 1.50 on the quality margin, and reviewer capacity bought back the most at +1.46. The quality floor decided whether a cliff showed up at all. It took a floor of 0.99 to turn most of the space fragile; at 0.95 there was none. SIGIL tested a prediction from the Accountability Rent working paper that review before commitment pays on irreversible decisions. The prediction failed in all 12 registered series. Give the reviewer a ledger of its own past damage and wrongful blocking at the top tier went from 30.1% to 42.7%. The lab published that as a refutation. AGENESIS-2 measured what a fleet of agents costs as it grows. Under a flat monitoring budget the cost exponent is 1.4665, with a 95% interval of 1.4101 to 1.5230. Give every agent its own check and it falls to 0.9001. Before publication the independent check turned up duplicate runs, which moved one verdict from supported to refuted. Trial Balance gave the auditor a ledger of what the agent owed. Reading only where that ledger failed to balance found 0.80 of skipped steps at 745 tokens. Log review topped out at 0.58, whatever the budget. AGENESIS , published in April, simulated first-mover adoption and found a premium of 210%, with the market tipping in 77% of markets. The product result The year's product is the open decision head, a small service that answers typed questions about a live case with a probability for each answer, on a model the operator runs. Its first version failed to beat a hand-written rule table, with a gap whose interval included zero. A retrained version, tested on 800 fresh cases, scored 0.728 against the rule table's 0.566, a gap of +0.161 with a 95% interval of 0.131 to 0.191. At first it wrongly closed cases more often than the project's bar allowed. The cause sat in the evidence pack the environment built, and Chapter 9 prints both readings. Corrections, and what comes next Chapter 11 lists every correction the pieces report, among them the duplicate runs in AGENESIS-2 and a retracted sentence in Error Independence. Schouw is scheduled for 13 October. The next Delegation Index reading freezes on 5 January 2027. 05 OCT 2026 Chapter 1 The Delegation Index What it measures The Delegation Index is one number a quarter, on a scale from 0 to 100. It sums up what the lab's own published campaigns say about how far enterprise work can go to AI agents under oversight when the oversight spend stays fixed. At 50, the cost of handing work over matches what oversight buys back at no extra spend. The construction was registered on 5 October 2026, before the Q3 numbers went through it. Only numbers printed in a published piece with a live public link on the freeze date enter. Anything internal stays out, drafts and unposted corrections included. Version 1, registered 5 October 2026 Freeze 5 October 2026, Chicago time Campaigns in 5 of 6 published Recompute independent, 6 October 2026, every value reproduced to two decimals The selection rule Each component reads the lab's best measured design in its campaign with oversight spend held fixed. The rule went on paper in advance so nobody could pick the flattering number afterward. Delegation price expresses one increment of delegation depth in increments of reviewer capacity, the largest positive lever in the Delegation Cliff space. Both coefficients sit on the same band-standardised scale, so the ratio has no units. Oversight yield uses plain review, the better of the two designs SIGIL measured before commitment, under the registered convex cost model. For audit detection the rule picks the checklist reader at a full read in Trial Balance, the best diary audit that has a pooled rate and an interval in print. Reviewer independence reads the five cross-family pairs in Error Independence, since that was the best pairing measured. Sprawl takes the fixed monitoring budget in AGENESIS-2. No other measured arm keeps spend flat while the fleet grows. The formula Every component is scored from 0 to 100 and clipped to that range. Each balance point comes straight from the component's own definition. Code Formula Scores 50 when P 100 × capacity ÷ (capacity + depth) one increment of capacity buys back exactly one increment of depth Y 50 × (1 + m), m the mean over tiers of value ÷ stake review breaks even A 100 × detection rate the audit finds as many skipped steps as it misses R 50 × (1 − φ) the second reviewer's errors are uncorrelated with the first's S 50 × (2 − β) oversight cost grows in proportion to the fleet The index is the plain mean of the five. Its band is the mean of the five lower ends and the mean of the five upper ends, which is the range the index would take if every component sat at the same end of its band at once. It has no coverage level, so the lab doesn't call it a confidence interval. Figure 1 The five components and the index, Q3 2026. Sprawl holds the reading down; the index band contains 50. Component scores on the 0 to 100 scale with their bands. Bands come from a published interval (A, S) or the published spread of seeds or pairs (Y, R); P has none. The index band is the mean of the band ends and carries no coverage level. Source: The Delegation Index, Q3 2026 reading, registered, publication pending; inputs from the five published pieces cited in Chapter 1. Rebuilt from the published numbers. The numbers as a table Component Score Band Delegation price, P 49.3 no band Oversight yield, Y 46.9 46.8 to 47.0 Audit detection, A 58.3 47.9 to 66.7 Reviewer independence, R 55.1 36.0 to 73.5 Sprawl, S 26.7 23.9 to 29.5 Delegation Index 47.2 40.8 to 53.2 The arithmetic Delegation price, P. "Pricing Agent Autonomy" prints a depth coefficient of −1.499 and a reviewer-capacity coefficient of +1.456, on the band-only Sobol fit, n = 4,948. P = 100 × 1.456 ÷ 2.955 = 49.27. One increment of depth costs 1.03 increments of reviewer capacity. The piece prints no interval, so P enters at its point value and adds no width to the band. Oversight yield, Y. The SIGIL paper prints the value of review against no oversight per tier and per seed. The pooled values run −0.13, −0.17, −0.18, −0.17 and −2.64 from tier 1 to tier 5. Dividing each seed mean by its stake, the square of the tier, and averaging over the five tiers gives m = −0.062, so Y = 46.90. Plain review lost about 6% of the stake per decision. The three seeds give a range of 46.78 to 47.01. Audit detection, A. Trial Balance prints omission detection at a full read with a per-step checklist of 0.583, 28 of 48, with a 90% cluster-bootstrap interval of 0.479 to 0.667. A = 58.3, with a band of 47.9 to 66.7. Reviewer independence, R. Error Independence prints the error correlation of each pair on the 47 contested items. The five cross-family pairs read +0.280, +0.013, −0.105, −0.227 and −0.469, with a mean of −0.102. R = 50 × 1.1016 = 55.08. The spread of the five pairs gives a range of 36.00 to 73.45, the widest band in the index. Sprawl, S. AGENESIS-2 prints an exponent of 1.4665 with a cluster-robust 95% interval of 1.4101 to 1.5230. S = 50 × (2 − 1.4665) = 26.68, with a band of 23.85 to 29.50. A higher exponent gives a lower score, so the band flips. The index. (49.27 + 46.90 + 58.30 + 55.08 + 26.68) ÷ 5 = 47.25, which rounds to 47.2. The lower ends average 40.76 and the upper ends 53.19. Behind the number Five of the lab's six published campaigns enter. Their scale is 139,646 units: 90,880 configurations, 30,000 supervised episodes, 18,360 audits, 150 generated invoices and 256 simulation units. The units behind the five quoted numbers come to 20,154. Those units are different kinds of object, so the total is only a count. The two thinnest legs, 48 checklist audits and 47 contested invoices, carry the two widest bands. Before Trial Balance entered, the other four components averaged 44.5, with a band of 39.0 to 49.8. Trial Balance went live on 1 October, inside the freeze window, and adding its audit component took the reading to 47.2. So all of that move comes from which campaigns are in. Sensitivities Oversight yield under the linear cost model scores 45.66, which moves the index to 47.0. Reviewer independence read from the same-weights pair, +0.447, would score 27.7, 1.0 above sprawl at 26.7. Audit detection read from the commitment-ledger audit at its lowest published cell, 0.92, would move the index to 54.0. That design has no pooled rate with an interval in print, so the registered rule leaves it out. It is printed here so readers can see it. Sprawl under per-agent monitoring, exponent 0.9001, would score 55.0. That arm raises spend with every agent, which breaks the fixed-spend condition. What stays out AGENESIS measures competitive dynamics between firms and prints no oversight quantity, so it stays out. Trial Balance's commitment-ledger audit is the strongest design the lab has measured; once a published piece prints a pooled rate with an interval, it enters through a new version. In SIGIL's after-the-fact audit arm, the paper itself calls the crossings post-hoc and unstable across seeds. The Delegation Cliff quality-floor shares describe the shape of the design space, and the price component already reads that campaign. What it does not claim The index reads one lab's synthetic measurements under the lab's own conditions. The Delegation Cliff coefficients inside it are descriptive slopes. SIGIL and Trial Balance each ran one model in the measured seat, so the reading holds for those models. The balance point at 50 marks where the lab's measured conditions cancel. Outside this index it means nothing. Rules that bind every reading Only numbers from live public pieces on the freeze date enter. The first paragraph of every release says the index reads the lab's own synthetic measurements. The headline never appears without its band and the count of campaigns and units behind it. Components with no band are named, and the release says the printed band is too narrow by their width. When the band contains 50, the release says the reading cannot tell balance from imbalance. The construction changes only through a new dated version, and the first release under it restates every earlier reading. Someone who did not build the reading recomputes it from the construction page and the cited pieces before release. A refuted result enters on the same terms as a supported one. SIGIL's refutation sits in oversight yield at full weight. The builder had read every input before fixing the balance points, because the inputs were public. The registration binds every reading after the first, which was made with the inputs already known. 06 OCT 2026 Chapter 2 How the lab works The protocol The lab studies how organisations hand work to AI agents and how they oversee it. Every campaign follows the same eight steps. If the lab tried to steer a result toward the answer it hoped for, the steps would leave a dated trace of the attempt. Ask one question. A campaign starts from a question a buyer or a regulator would recognise, such as whether a second model checks the first or what a liability ledger does to a reviewer. It names the quantity that would answer the question and the null result that would mean the idea adds nothing. Register the prediction, the bands and the kill conditions. Before any unit runs, the lab writes down the hypotheses and the band each number has to land in to count. The statistics and the conditions that stop the run go into the same document, which is dated and frozen. A later change goes in as a dated amendment, written before anyone reads the data it touches. A small calibration stage first checks that the task is neither trivial nor impossible. That gate checks the instrument works and nothing more. The lab learned where that line sits by halting a campaign whose gate had demanded the result it existed to measure. Attack the design. A reviewer whose job is to break the headline goes over the design and the analysis. Confounds and statistics that hide a split are the usual targets, along with ceilings and scoring rules that favour one arm. Whatever the review finds gets fixed before the run or stated as a limit in the piece. Run to the registered rule. The campaign runs on open models on local hardware the lab owns, so anyone with the code can repeat a run. It stops where the registration says it stops. A failed gate or a spent compute budget ends the run, and the lab analyses whatever is complete at that point. A run going badly keeps its length and its thresholds. Verify from the raw files. A separate pass with fresh code and no access to the original analysis recomputes the reported numbers from the raw records. Each piece says how many numbers were checked and how many matched. Where the two passes disagree, someone chases the gap down, and the piece carries the verified value. Publish the negative results. A refuted hypothesis appears with the same care as a confirmed one, next to the prediction it contradicted. A null result is reported as a bound. Each piece keeps what was measured apart from what the lab infers. An explanation stays labelled as one until a test has run against it. Release the code. Each campaign keeps its registration, harness, unmodified ledger and analysis together, so a reader can trace every published number to the rows behind it. The code and records go out with the pieces as each campaign is cleared for release. Log every correction. When a published number turns out wrong, the correction goes up in the open with a date, and the original stays visible beside it. The measured claims index in Appendix A keeps a row for every headline number, so a correction has a fixed place to land. What the steps caught this year During 2026 the verification pass on AGENESIS-2 found 95 duplicate pairs of simulations and moved a verdict from supported to refuted before the paper went out. Trial Balance's first attempt halted at its calibration gate with zero production units, and the gate rule was rewritten the same evening. SIGIL's registered prediction failed in every series, and it went out as a refutation. Chapter 11 lists every correction the pieces report. The house format The campaign chapters that follow share one layout. The question comes first, then what was registered before any unit ran, then what was measured and at what scale. The finding carries its number and, where the piece prints one, its interval. A chart rebuilt from the published numbers follows. Each chapter closes with what the campaign does not claim and with where the numbers were checked. 04 APR 2026 Chapter 3 AGENESIS: first-mover adoption Published 4 April 2026, Medium Scale 100 simulation runs, 4 industry verticals Verdict measured Pre-registration not recorded The question What happens to an industry when one company adopts AI agents before its competitors do? AGENESIS was the lab's first published campaign, and it asked that question of a simulated market. What was registered The site records no pre-registration for AGENESIS and no verification status. The findings are in the piece. This chapter reports them with no registration line, since none exists on the record. What was measured AGENESIS is an agent-based simulation of enterprise AI adoption. The published article reports 100 independent simulation runs across 4 industry verticals, under 5 adoption configurations and 3 macroeconomic conditions. The finding 210% First-mover premium across 100 runs. The market tipped in 77% of markets. The firm that adopted first earned a premium of 210% over its competitors, and the market tipped toward one firm in 77% of the simulated markets. The lab's later campaign on sprawl cost reused the same simulator and added an agent lifecycle to it. What it does not claim AGENESIS has no agent lifecycle. Agents are created once and never retired, and the AI cost a firm pays has nothing to do with how many agents it runs. A firm with fifteen agents and a firm with a hundred and fifty pay the same at the same level of adoption. The sprawl-cost campaign in Chapter 7 was built to fill that gap. AGENESIS prints no interval and no oversight quantity, so the Delegation Index leaves it out. Where it was verified The site records no verification status for this piece, and the two numbers above are quoted from the published article. 02 SEP 2026 Chapter 4 Error Independence Published 2 September 2026, Medium Scale 150 items, four model configurations Verdict measured Ground truth generated with each item The question A common resilience design runs the same model a second time, on different accelerators or in a second region, and treats the second answer as a check on the first. Civil aviation learned long ago that two identical computers fail the same way at the same moment. Flight control systems run on processors from different makers, with software written by teams kept apart, and the industry calls that dissimilar redundancy. This study asked whether a second copy of a model, on different hardware, makes different mistakes. What was registered Ground truth is generated with each item, so the correct answer is known before any model sees it. No model grades another, and so there is no judge inside the measurement. The site records no pre-registration beyond that design. For each pair of models the study computed the phi coefficient on error indicators. Phi sits near zero when two models go wrong on different items and near one when they go wrong together. What was measured Four model configurations scored the same 150 invoice-extraction items. The key comparison ran model A twice, in two builds with identical weights. The builds differed in quantisation and in the silicon and serving stack beneath them. Models B and C came from other families. 103 of the 150 items were unanimous. On 54 of them all four models were right. All four missed the other 49. Those items measure how hard an item is, so the headline uses the 47 items where the models disagreed. The finding Before conditioning, every pair looked correlated, from 0.556 to 0.827. On that reading you would give up on a second reviewer altogether. +0.447 Error correlation, identical weights on different silicon, 47 contested items. Cross-family pairs averaged −0.102. On the contested items, the pair with identical weights stayed at +0.447, the highest in the matrix. The five pairs from different model families averaged −0.102, and model A's first build against model B sat at −0.469. Changes of machine and serving setup left the correlation standing. Only a change of model family moved it through zero. Figure 2 Pairwise error correlation before and after conditioning on item difficulty. Only the pair with identical weights stays high. Pairwise error correlation (phi) on all 150 items (open ring) and on the 47 contested items (filled dot). The highlighted pair runs identical weights on different hardware. Source: “When the Second Opinion Shares the Blind Spot”, Medium, 2 September 2026. Rebuilt from the published numbers. The numbers as a table all 150 items 47 contested items Same weights, two machines 0.827 0.447 Model A build 2 + model C 0.810 0.280 Model A build 1 + model C 0.729 0.013 Model B + model C 0.617 −0.105 Model A build 2 + model B 0.607 −0.227 Model A build 1 + model B 0.556 −0.469 What it does not claim The same-family reading comes from one pair, the only two systems in the set with identical weights. One pair counts as an observation, and it would take more pairs before anyone quotes it as a rate. The five cross-family pairs run from −0.469 to +0.280. A mean of −0.102 sits inside a spread that wide and describes it poorly. Two families can still fail together; the reading covers these families on this task. The study ran one task, chosen because its ground truth cannot be argued with. Open-ended work is untested. The article also reported a model-free checker as a floor for independence. A later note retracted that sentence. Chapter 11 gives the details. Where it was verified Model scoring ran through a separate verification path, and every model-pair correlation survived the later review that retracted the checker sentence. 03 SEP 2026 Chapter 5 Delegation Cliff Published 3 September 2026, Medium, as "Pricing Agent Autonomy" Scale 90,880 configurations, zero error units Verdict mixed Certificate a 200-pair certificate passed, 95th percentile 0.0117 The question Any organisation deploying AI agents is betting that pushing work further from human review costs less than keeping it close, and few of them put a price on that bet. This campaign tried to. Where does delegation stop paying, and which design choices buy the margin back? What was registered Four predictions went on the record before the sweep. H1. The design space has a cliff, a sharp boundary between robust and fragile configurations. H2. Better observability outranks raw reviewer capacity. H3. Agent self-check is the dominant lever. H4. An audit death spiral forms, where eroding trust raises audit load until capacity saturates and quality falls. What was measured The campaign simulated an agent organisation with twelve design dimensions, and this chapter turns on two of them, delegation depth and reviewer capacity. The sweep scored 90,880 configurations on the margin between the quality each organisation produces and the floor it must hold. It finished with zero error units. A revalidation certificate reran 200 random configurations from scratch: the median difference was 0.002 and the 95th percentile 0.0117, against a tolerance of 0.05. The finding −1.50 Change in quality margin per added level of delegation depth. Reviewer capacity buys back +1.46. Delegation depth carries the largest weight in the space, with a negative sign: each increment costs 1.50 on the margin. Reviewer capacity buys back the most at +1.46. Model capability at +0.95 and self-check calibration at +0.92 follow, close enough to read as tied. Task coupling is the second-largest cost, at −0.96. Figure 3 Linear-probe coefficients on the delegation margin. Delegation depth costs the most; reviewer capacity buys back the most. What moves the delegation margin: linear-probe coefficients on the Sobol sample, band-only, n = 4,948. Positive buys margin, negative spends it. Source: “Pricing Agent Autonomy”, Medium, 3 September 2026. Rebuilt from the published numbers. The numbers as a table Design dimension Coefficient Delegation depth −1.499 Reviewer capacity 1.456 Task coupling −0.958 Model capability 0.949 Self-check calibration 0.915 Rework cost −0.803 Verification depth 0.413 Verification coverage 0.308 Detection lag −0.259 Workload volatility −0.147 Trust response 0.102 Escalation latency −0.007 The quality floor decides the failure geometry. At a 0.95 floor, 60% of the space is a broad transition band with no cliff. Raise the floor to 0.99 and the cliff appears, with 79% of the space fragile. refuted H1 at the 0.95 floor. H2, H3 and H4 not supported. The observability prediction failed. Capacity sits at +1.46, verification coverage at +0.31 and detection lag at −0.26. Interaction terms between observability and capacity came out real and positive, which makes observability a complement to reviewer capacity, and the hypothesis fails all the same. As for the spiral, mean margin rises across trust-response quintiles, even inside the high-stress stratum where depth and coupling are both high, so it never formed. What it does not claim The coefficients are descriptive slopes on the Sobol sample inside the transition band, n = 4,948. They describe how the margin moves where a boundary exists, and they carry no causal weight over the whole space. The 0.99 result rests on one 256-unit calibration sweep. The 0.95 result is replicated on the 8,192-unit Sobol set. The article gives the two legs different weight for that reason. Reviewer capacity is a parameter in this design. Reviewers here never tire. Capability and self-check differ by 0.034. The article reads them as tied. Where it was verified Every result file carries a finite margin and a null error code, which is how the zero-error count was checked. The 200-pair revalidation certificate passed. The published coefficients come from the corrected band-only fit; Chapter 11 describes the correction that produced them. 10 SEP 2026 Chapter 6 SIGIL Published 10 September 2026, Medium, as "Accountability Makes Oversight Worse" Scale 30,000 supervised episodes Verdict refuted Recompute 0 mismatches in 29,902 scored units The question Once a reviewer holds the power to stop an AI worker, the usual next step is to make it accountable for what it lets through. This campaign asked whether an AI supervisor authorises differently when it carries a running record of the damage its own past approvals caused. It also asked how the value of its review changes as decisions get harder to undo. What was registered The hypothesis came from the Accountability Rent working paper. It predicted that review before commitment loses value on reversible decisions, turns positive past some level of irreversibility, and climbs steeply from there. A flat response falsifies it, as does a concave one or the absence of any crossing. The output contract was frozen on 30 August 2026, before any harness code existed. It fixed the scoring and a gate that required the unsupervised worker to err between 10% and 35% on a pilot. It also made a rerun under a second, linear cost model mandatory. One amendment, recorded before the second pilot, changed only the task generator and harness capacity and pre-committed a stopping rule against retuning more than once. What was measured An AI worker approved or rejected synthetic firm decisions across five tiers, from invoice approvals that reverse next cycle to public statements that cannot be taken back. The same base model sat in a second seat and reviewed it under four designs: no oversight; an after-the-fact audit of a 20% sample; review with the power to block before commitment; the same blocking review carrying a liability ledger of the harm its own past approvals caused. Worker and reviewer share one model on purpose, so any difference comes from the incentive design. Every decision was scored against programmatic ground truth. The campaign ran 30,000 episodes on local hardware, in two full campaigns, one per cost model, with three seeds per cell. The finding refuted The registered prediction failed in all 12 series, under both cost models. Review never crossed zero, in any of the 12 registered series. Its value collapsed at the most irreversible tier, where the theory wanted it to rise, to −2.64 per decision against no oversight. The liability ledger made it worse, at −3.81. 42.7% Tier-5 decisions wrongfully blocked with a liability ledger. Plain review blocked 30.1%. At that tier plain review wrongly blocked 30.1% of decisions, and the ledger pushed the share to 42.7%, since refusing was the one action it never charged for. Figure 4 Net value of blocking review against no oversight by irreversibility tier. Both designs stay below zero and fall at tier 5; the liability ledger falls further. Net value per decision against no oversight, by irreversibility tier, convex damage model, pooled over three seeds. Below zero, oversight costs more than it saves. Source: “Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers”, Medium, 10 September 2026. Rebuilt from the published numbers. The numbers as a table Irreversibility tier plain review liability ledger 1 −0.13 −0.14 2 −0.17 −0.20 3 −0.18 −0.16 4 −0.17 −0.51 5 −2.64 −3.81 A reviewer that can only block can fix one kind of mistake, a wrongful approval. The share of worker errors that were wrongful approvals fell from 93% at tier 1 to 1.4% at tier 5, one false approval against 68 false rejections. What it does not claim This is a controlled simulation with one open-weight model family in both seats, scored against constructed ground truth. The liability ledger is a thin, simulated form of accountability. A follow-up registered to test stronger models stopped before it ran an oversight trial. Four other models got between 97.3% and 100% of cases right unsupervised, which left a reviewer nothing to catch, so the result holds for this setting only. The after-the-fact audit arm shows positive cells in some series. Its seeds disagree on where and whether it crosses zero, and the paper reports those cells so nobody has to dig for them. The contract-clause tier produced no worker errors across both campaigns, so its value figures are pure oversight cost. Where it was verified A separate verifier recomputed every table from the raw episode ledgers and found zero mismatches on 14,950 convex and 14,952 linear scored units, 29,902 in all. It also confirmed that the grid was complete and that no review unit carried ledger state. That pass surfaced one documentation defect in the tail pre-registration, recorded as erratum E1 and listed in Chapter 11. 17 SEP 2026 Chapter 7 AGENESIS-2: sprawl cost Published 17 September 2026, Medium Scale 256 simulation units on held-out seeds Verdict mixed Correction 95 duplicate pairs found before publication The question The governance literature says an agent portfolio you cannot retire from fast enough costs more than proportionally as it grows, though nobody had put a number on it. This campaign asked how fast the cost of a simulated agent fleet grows with its size, and whether the answer depends on how the monitoring budget is set. What was registered Seven hypotheses with frozen thresholds went on the record before the two seeds that count had ever run. The kill condition came from the review that commissioned the work: if cost turned out linear in agent count, the design was falsified and the project would retire. Hypothesis Registered criterion H1, superlinear under a flat budget pooled 95% lower bound above 1.10 H2, clean accounting where nothing binds exponent inside 0.85 to 1.10 on the per-agent arms H4, the exponent flattens at large sizes intervals on the four smallest and four largest sizes do not touch H5, overlap cuts both ways mask falls and detection lag rises H6, four factors earn their place each non-detection factor moves the exponent by at least 0.10 H7, detection and review are serialised capacity moves the exponent by less than 0.05 under a flat budget H3, an ordering of the four factors, was registered as descriptive only. What was measured In the simulation, agents ramp up and then drift silently out of spec, doing damage while nobody notices. A flagged agent queues for review and gets retired. Cost is the harm from degraded agents still running plus staff time spent reviewing. No cost term depends on how many agents exist, so any growth beyond linear has to come through the review queue and the detection lag, or through overlap between agents. 256 runs covered eight portfolio sizes from 57 to 787 agents, over an 18-month horizon in weekly ticks. The finding 1.4665 Sprawl cost exponent in fleet size, monitoring budget held flat. 95% interval 1.4101 to 1.5230. Under a flat monitoring budget, cost grows as fleet size to the power 1.4665, with a cluster-robust 95% interval of 1.4101 to 1.5230. Doubling the fleet multiplies cost by about 2.76. Where every agent carries its own check and review grows with the fleet, the exponent is 0.9001, with an interval of 0.8664 to 0.9337, about 1.87 per doubling. You can see the mechanism in the detection lag. With a fixed inspection budget, the time from an agent going bad to someone noticing runs from 2.78 weeks at 57 agents to 22.49 weeks at 787. Under the flat budget, changing review capacity produced a bit-identical simulation in 55 of 64 matched cases, because the review queue never formed, which puts detection upstream of review. refuted H4, the regime ceiling, after the duplicate correction. H6 refuted as well. The independent check found 95 duplicate pairs among the 256 rows, 161 distinct simulations in all. Correcting for them left every point estimate in place and widened the intervals. That flipped H4 from supported to refuted: fitted on the four smallest sizes the exponent is 1.4949, on the four largest 1.3238, and the corrected intervals overlap by 0.0145. H6 failed too. Retirement policy moved the exponent by 0.0622 and overlap by 0.0892, both under the frozen 0.10, so the follow-on study will carry two factors. Figure 5 The regime-ceiling test before and after the duplicate correction. The registered intervals clear each other; the corrected intervals overlap. Cost exponent fitted on the four smallest and the four largest portfolio sizes under diluted monitoring, with 95% intervals as first reported and after clustering on the result signature. The frozen test needed the intervals to stay apart. Source: “Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics”, Medium, 17 September 2026. Rebuilt from the published numbers. The numbers as a table Fit Exponent 95% interval Smallest four, as reported 1.495 1.415 to 1.574 Largest four, as reported 1.324 1.262 to 1.386 Smallest four, corrected 1.495 1.404 to 1.586 Largest four, corrected 1.324 1.229 to 1.418 What it does not claim This is a calibration on one simulated industry, one company size, one degradation rate and one 18-month window. It was registered as a calibration that cannot confirm anything on its own. The value 1.47 awaits a confirmatory study, and until then the lab leans on the shape of the curve. The flattening at large sizes is an observed direction with a failed test. Where it was verified An independent agent re-derived every number from the raw ledger without reading the analysis code. It found that only 161 of the 256 result signatures were distinct. The campaign's own guard had checked replicate distinctness inside each cell, and every duplicate pair sat across two cells. The paper prints the registered and corrected numbers together, and Chapter 11 sets out the correction. 01 OCT 2026 Chapter 8 Trial Balance Published 1 October 2026, Medium Scale 18,360 audits on one auditor model Verdict supported Rescore point estimates equal to the registered analysis The question An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds the gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. This campaign asked whether the same idea works for agents. If the layer that hands out the work writes down what the agent owes, and the harness writes down what the agent ran, can an auditor find skipped steps for a fraction of the reading cost by looking only where the two books fail to balance? What was registered The registration predicted that imbalance-first audit would reach 0.80 omission detection at a third of log review's reading cost or less, at 8, 12 and 16 requisitions. A balanced book, it said, would hide a wrong amount carried on both sides. And a reader of the log would miss a skipped step far more often when the agent wrote a line claiming the step was done. The first attempt halted at its calibration gate on 11 September with zero production units. Its gate had required a reading to move that the campaign existed to measure, so a strong true result would have failed it. The lab rewrote its gate rule that evening. The second attempt was registered and frozen before its first unit, with one added arm, a log reader told to tick off every step the procedure requires before it gives a verdict. What was measured A procurement workflow with seven steps ran at seven lengths, from one requisition to sixteen. Five conditions were injected into each diary: no fault, a skipped step, a skipped step with a line claiming it was done, a wrong amount, and a wrong amount carried on both sides of the book. Four auditors read each diary under reading budgets from 200 to 12,800 tokens, seven for most arms. Log review reads the diary top down. The imbalance-first auditor opens the trial-balance report and goes on to read only the requisition that failed to balance. A random sample of the same size controls for the effect of reading less. The checklist reader gets the log reviewer's diary plus an instruction to mark every owed step present or absent. The auditor was one open model under a fixed reply schema, on local hardware. It ran 18,360 audits, including 3,240 on diaries written by a different model running the workflow for real. The finding Imbalance-first reached 0.80 omission detection at the smallest budget, 745 prompt tokens at eight requisitions, and sat between 0.92 and 1.00 in every cell. Log review stayed under 0.80 at every budget and length, peaking at 0.58. Per thousand tokens read, imbalance-first produced about 21 times as many correct omission findings as log review at its best budget, 1.34 against 0.063. 745 Prompt tokens at which imbalance-first reaches 0.80 omission detection, eight requisitions. 90% interval 740 to 749. Figure 6 Detection at a full read by fault class. Imbalance-first finds every fault that unbalances the book and none of the consistent commissions. Share of injected faults named correctly at a full read, by fault class, synthetic primary seeds, 84 per class. Source: “A Trial Balance for Agent Omissions”, Medium, 1 October 2026. Rebuilt from the published numbers. The numbers as a table Fault class Imbalance-first Checklist reader Log review Silent omission 1.00 0.73 0.48 Phantom completion 1.00 0.36 0.18 Value commission 1.00 0.98 0.98 Consistent commission 0.00 0.95 1.00 The balance check has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, so the trial balance is clean, and the check found 0 of 1,044 of them. A full read of the diary finds those almost every time, because the vendor's quote still shows the true amount. The cheapest result in the campaign came from the checklist instruction. With the whole diary in front of it, plain log review found skipped steps at 0.25, with a 90% interval of 0.19 to 0.31. The same auditor told to tick off each owed step found 0.58, with an interval of 0.48 to 0.67, and accused no clean diary at a full read on the primary seeds. What it does not claim One auditor model ran. The second registered family was excluded before its calibration gate, and the lab claims nothing about it. Log review never reached the bar, so the registered cost ratio is unbounded and the 21-times figure is the finite comparison to quote. The campaign cannot split the checklist's gain between the list and its definition of evidence. One workflow, one fault per diary, and real diaries only up to three requisitions. Nobody priced the cost of building a commissioning layer that writes obligations in ledger form. Where it was verified An independent rescoring from the raw ledgers matched every point estimate, with 90% cluster-bootstrap bounds differing only by bootstrap noise. 06 OCT 2026 Chapter 9 The open decision head Status executive briefing published 7 October 2026; paper and blog follow Scale 800 + 200 onboarding cases, 800 fresh cases, 400 + 200 cases on a second process Verdict mixed Rescore every phase, from the raw case files The question Most decisions inside a governed business process are questions with consequences. Does this case have the evidence it needs? Should someone ask the supplier for a document, or think harder about the one on file? A decision head answers questions of that shape. You hand it the state of a case and a list of typed questions, yes or no or pick one of these, and it returns a probability for each answer, a number a workflow can threshold and audit. Jev, a hosted service from TypeSafe AI, showed the shape. The lab built an open version that applies the same request and response pattern on an open-weights model running on hardware the operator owns. The project asked two things. Can a controller tell "this case needs a document nobody has" from "this case needs someone to interpret a document already on file"? And does an open, self-hosted head asked the same typed questions get there? How it works The head renders the state of a case once, as a shared prefix. Each question then goes in its own call, appended to that prefix, so no question's prompt contains another question's text. Each call asks for exactly one token at temperature zero and reads the model's token probabilities over the answers the question accepts. If none of the accepted answers appear, the head returns a flagged uniform distribution, so it never invents a number. Every decision writes an audit row with the question and its probabilities, raw and calibrated. The row also carries the latency and a hash of the exact request. What was registered Five hypotheses with numeric predictions went on the record before any evaluation data existed. The first, discrimination, predicted the head's decision-class accuracy between 0.64 and 0.76 and required it to beat the rule table and the typed-sample arm with a case-clustered bootstrap interval excluding zero. Three adversarial reviews cleared the design, and the first two returned "must not run". A later registration set 0.04 as the wrong-claim rate that would trigger a redesign of the claim guard, and 0.02 as the bar a shipped head must meet. What was measured The controller sits inside a synthetic supplier-onboarding workflow and picks the next of six approved actions: continue, retrieve a document, request one, consult an expert model, hand the case to a person, or claim the case complete. Every case is generated from templates with no model involved, so ground truth is correct by construction and no client data appears anywhere. The first round used 800 evaluation cases and a 200-case challenge set with contradictory evidence. A second registration retrained the head and tested it on 800 fresh cases. A third moved everything to a second process, accounts-payable exception routing, on 400 and 200 cases. All of it ran on local hardware. The first round refuted H1, discrimination, on the first registered split. The zero-shot head did not beat the rule table. On 800 evaluation cases the zero-shot head scored 0.545 against the rule table's 0.575. The gap of −0.029 has a 95% interval of −0.060 to +0.001, and H1 was rejected. Completion at cost was rejected as well: the head completed fewer cases correctly and called the expert model about a hundred times as often as the rule table, 0.376 calls per case against 0.004. Its decision-class calibration error came in at 0.5247, more than six times the registered ceiling of 0.08. The second round 0.728 Retrained head, decision-class accuracy on 800 fresh cases. Rule table 0.566. Gap +0.161, 95% interval 0.131 to 0.191. The retrained head scored 0.728 on the 800 fresh cases, against 0.566 for the rule table, a gap of +0.161 with a 95% interval of 0.131 to 0.191. On the challenge set it scored 0.739 against 0.565, a gap of +0.171 with an interval of 0.115 to 0.224. A separately fine-tuned learned head scored 0.701 on the fresh cases. Its gap to the retrained head, −0.027 with an interval of −0.056 to +0.002, fails the project's own acceptance rule, so the learned head ships as experimental and the retrained head stays the default. Figure 7 The head's gap to the rule table, run by run. Only the retrained head on the first process clears zero. Decision-class accuracy of the decision head minus the hand-written rule table on the same cases, with case-clustered paired bootstrap 95% intervals. Right of zero, the head picks the right kind of next step more often. Source: “The Open Decision Head: An Executive Briefing”, LinkedIn, 7 October 2026; the paper and blog post are published on dinand.com . Rebuilt from the published numbers. The numbers as a table Run Head minus rule table 95% interval First round, zero-shot head, 800 cases −0.029 −0.060 to 0.001 Retrained head, 800 fresh cases 0.161 0.131 to 0.191 Retrained head, 200-case challenge set 0.171 0.115 to 0.224 Second process, 400 cases 0.002 −0.029 to 0.034 Second process, 200-case challenge set 0.022 −0.023 to 0.066 Wrong claims and the evidence pack On the fresh cases, the retrained head wrongly claimed a case complete at a rate of 0.0466, above the 0.04 trigger. A taxonomy of its 37 wrong claims, checked independently, found the cause in the environment. In 31 of them the pack builder had padded the evidence pack with extra supporting documents. Of the 101 extra documents, 91 fail the harness's own validity check for entity and currency. The fix was registered before it was computed: the pack builder now checks every supporting document with the same predicate the ledger already used. Figure 8 Wrong-claim rate before and after the pack-builder fix. Both heads clear the 0.02 bar only on the validated pack. Share of the 800 fresh cases closed with a claim the registered exact-match scorer marks wrong, before and after the pack builder checks every supporting document for entity and currency. Source: “The Open Decision Head: An Executive Briefing”, LinkedIn, 7 October 2026; the paper and blog post are published on dinand.com . Rebuilt from the published numbers. The numbers as a table original pack validated pack Rule table 0.0112 0.0088 Retrained decision head 0.0466 0.0113 Learned head 0.0537 0.0163 After the fix the wrong-claim rate falls to 0.0113 for the retrained head, 0.0088 for the rule table and 0.0163 for the learned head. Every case that changed moved from wrong to verified, and no other case moved. Both heads meet the 0.02 bar on the validated pack and miss it on the original pack. Nobody found the fix until the trigger had fired and its cause was traced. Retiring the claim guard Three guards were tried against wrong claims, and all three were retired. Voting across several readings of the pack separated wrong claims from right ones at an AUC of 0.56, below its registered gate. The deterministic evidence guard first reported 79 of 79 caught, until a review found its checks reused the functions that define a wrong claim; the lab withdrew the number as circular. Opening the cited document, a reading layer reached an AUC of 0.52 on its own and added nothing on top of the deterministic check. The validity check now lives inside the pack builder and calls no model. A second process On accounts-payable routing the head scored 0.650 and the rule table 0.648. The gap of +0.002, with an interval of −0.029 to +0.034, missed the registered bar, and the head's lead from the first process was gone. A learned head trained on onboarding scored 0.522 when pointed at the new process with no retraining. Retrained on 100 labelled cases from the new process, 271 decisions, it scored 0.663 and came level with the head. A version with every input from the judging head removed scored 0.660, which closed the question of whether the learned head was leaning on the head's own answers. Consistency and calibration In the first round the open head changed its answer on 4.95% of repeated decisions, while a second campaign shared its endpoint. Re-measured with nothing else attached and one call in flight, it changed on 0.05% of repeats on both processes. At four calls in flight the figure rose to 3.40% and 1.95%, with the disagreements clustered on one worker thread, the mark of calls batched together. The shipped service now defaults to one call in flight, which costs throughput and leaves accuracy where it was. At the item level, the retrained head's probabilities were already calibrated: expected calibration error 0.0178 before any refit and 0.0143 after one. Two open encoders released in September, encoder A and encoder B, ran on the same benchmark. Fine-tuned on 1,567 labelled onboarding decisions, encoder A scored 0.815 on the challenge set, above the retrained head. With only 201 decisions from the second process to learn from, it fell below the rule table there. What the head offers The open head competes on accuracy and on who owns the data and the tuning. A live case stays on the operator's network. Because the calibration fit runs on the operator's own labelled cases, it produces a reliability table an auditor can inspect, and each decision leaves an audit row behind it. To change behaviour, the operator relabels data. On speed, the hosted service answers a question in 0.35 seconds; the open head's fast mode took a median of 3.1 seconds under load on local hardware. A workflow gate fires about 3.4 times per case, inside steps that take minutes to days, so for most gates accuracy decides it. The code is public under Apache-2.0 at https://github.com/dtinholt/decision-head. What it does not claim Every case is synthetic, from one generator family, and the training labels and the scorer share one oracle. Real supplier files are messier. The second process shares the first one's blocker mix by construction. A process built differently is registered as future work. The retrained head clears its wrong-claim bar only on the validated pack, and the fix came after the trigger fired. The expert-consultation phase and its blind adjudication sample have not run. The head and the worker ran on one quantised base model, and nothing here measures another. Where it was verified Every phase was rescored from the raw case files by a party that did not build the scorer, importing the scorer's own metric functions. The consistency and abstention re-measurements matched on 28 of 28 cells within a tolerance of 0.0005. The encoder runs were recomputed cell by cell within the same tolerance and audited for leakage and training overlap before anyone read a hypothesis. 06 OCT 2026 Chapter 10 What comes next Forthcoming Schouw is scheduled for 13 October. Its numbers stay out of this report until the piece is live, and it enters the Delegation Index at the first freeze after that date if it reads one of the five registered quantities. The index The next Delegation Index reading freezes on 5 January 2027. Campaigns published by then enter under the rules in Chapter 1. A campaign that reads a quantity none of the five components covers enters only through a new dated version, and the release that first uses it restates the Q3 2026 reading beside the original. The decision head The decision-head code is public under Apache-2.0 at https://github.com/dtinholt/decision-head (release v0.1.0, 7 October 2026). A transfer test on a process with a different blocker mix and obligation structure is registered, along with the expert-consultation phase that has not yet run. 06 OCT 2026 Chapter 11 Corrections logged Why this page exists The lab publishes its corrections, and the original number stays visible beside each one. Every entry below is reported in the piece it corrects or in a note attached to it. Delegation Cliff: censoring bounds treated as measurements Corrected before publication on 3 September 2026, in "Pricing Agent Autonomy". The published coefficients come from the Sobol sample restricted to configurations inside the transition band, n = 4,948, "with no censoring bounds treated as measurements," as the article puts it. An earlier fit over all units had averaged the censoring bounds as though they were measured margins. That pulled every slope toward zero and ranked self-check ahead of model capability. On the corrected fit, reviewer capacity leads at +1.46, with capability at +0.95 and self-check at +0.92, read as tied. No conclusion of the campaign changed, and every figure in the article is from the corrected fit. AGENESIS-2: duplicate simulations Corrected 8 September 2026, before publication on 17 September, in the paper "Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget". An independent agent re-derived every number from the raw ledger and found 95 pairs of duplicate simulations among the 256 banked rows, 161 distinct simulations in all. The registered guard had checked distinctness inside a cell, and each pair sat across two cells. Treating the duplicates as independent had made the standard errors about 36% too small on the primary pool. No threshold moved. Clustering the standard errors on the result signature left every point estimate in place and flipped H4 from supported to refuted: the intervals that had cleared each other by 0.0297 now overlap by 0.0145. H7 was restated in the form the ledger supports directly, 55 of 64 matched cases bit-identical. The paper shows the registered and corrected numbers side by side. Error Independence: the model-free checker Correction note to "When the Second Opinion Shares the Blind Spot", dated 8 September 2026. The article reported a rule-based checker, running no model, that landed "at −0.155 to −0.012 against every model" and served as a floor for independence. The checker had a bug. Its pattern for the total field matched "Subtotal" first, so it was testing whether subtotal plus tax equals subtotal. A first repair produced a second wrong reading, traced to a defect in the test documents, where scan-style character swaps left some demanded strings absent from the document. The note retracts both the original sentence and the first repair. The model-pair correlations stand, because model scoring ran through a separate path, and the gap between same-weights and cross-family pairs widens once the defective items are removed. SIGIL: erratum E1 Reported in the paper "Accountability Makes Oversight Worse", 10 September 2026. The tail pre-registration's loss table left out that the scoring applies the 0.60 audit-recovery factor to every audited error, false rejects included. The analysis used the rule the campaign ran, which is why the recompute found zero mismatches. The omission touches only the audit condition, and no verdict rested on it. Read literally, the table's rule would have produced 48 mismatches in the convex campaign and 51 in the linear one, which is how the verifier found it. Trial Balance: the calibration gate Reported in the paper "A Trial Balance for Agent Omissions", 1 October 2026. The first attempt halted at its calibration gate on 11 September with zero production units. One gate clause required the reading budget to move a result the campaign existed to measure, so a strong true result would have failed it. The rule was rewritten that evening. The second attempt was registered and frozen before its first unit. Decision head: scorer defect, circular guard, evidence pack Reported in the decision-head paper . On the challenge set the first-round arms first appeared to claim a case complete wrongly in 21% to 28% of cases. Error analysis traced most wrong claims to the scorer, which marked a pack wrong for citing a second, equally valid document. Under a predicate-equivalent reading the rates across the arms in the paper's table sit between 1.5% and 9.5%. A deterministic claim guard's 79-of-79 catch rate was withdrawn as circular. The wrong-claim fix in the pack builder was found after the registered trigger fired, and both readings are published. 06 OCT 2026 Appendix A Measured claims index How to read it One row for every headline number. Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Intervals are 95% unless the piece says otherwise; Trial Balance intervals are 90% cluster bootstrap. Campaign What was measured Number Interval Source date Trial Balance Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions 745 740 to 749 1 October 2026 Best omission detection by plain log review at any budget or trace length (0.80 never reached) 0.58 1 October 2026 Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitions about 21 times (1.34 against 0.063) 1 October 2026 Omission detection by plain log review reading the whole diary 0.25 0.19 to 0.31 1 October 2026 Omission detection reading the whole diary with a per-step checklist instruction 0.58 0.48 to 0.67 1 October 2026 Wrong amounts carried on both sides of the book found by the balance check 0 of 1,044 1 October 2026 Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179) 0.798 0.738 to 0.857 1 October 2026 Omission detection, imbalance-first over a random sample of the same reading budget +0.907 0.891 to 0.925 1 October 2026 AGENESIS-2 Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust) 1.4665 1.4101 to 1.5230 17 September 2026 Sprawl cost exponent in fleet size, per-agent monitoring with review that scales 0.9001 0.8664 to 0.9337 17 September 2026 Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring 2.76× against 1.87× 17 September 2026 Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical 55 of 64 17 September 2026 Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents 2.78 to 22.49 17 September 2026 Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145) 1.4949 against 1.3238 1.4037 to 1.5862; 1.2294 to 1.4182 17 September 2026 Banked runs the independent verification found to be duplicates 95 duplicate pairs (190 of 256 rows; 161 distinct simulations) 17 September 2026 SIGIL Registered series showing the predicted zero crossing in the value of review 0 of 12 10 September 2026 Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model) −3.81 against −2.64 10 September 2026 Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger 30.1% against 42.7% 10 September 2026 Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 5 93% to 1.4% 10 September 2026 Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived) 0 of 14,950 convex and 0 of 14,952 linear 10 September 2026 Delegation Cliff Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948) −1.50 3 September 2026 Change in quality margin per increment of reviewer capacity, the largest positive lever +1.46 3 September 2026 Model capability against agent self-check calibration (read as tied) +0.95 against +0.92 3 September 2026 Share of the design space in a broad transition band at a 0.95 quality floor (no cliff) 60% 3 September 2026 Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep) 79% 3 September 2026 Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.05 0.0117 3 September 2026 Error Independence Error correlation (phi), identical weights on different silicon, contested items only +0.447 2 September 2026 Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280) −0.102 2 September 2026 Error correlation (phi) across all pairs before conditioning on item difficulty 0.556 to 0.827 2 September 2026 AGENESIS First-mover premium across 100 runs; share of markets that tipped 210%; 77% 4 April 2026 Decision head Zero-shot head minus rule table, decision-class accuracy, first registered split, 800 cases −0.029 −0.060 to +0.001 7 October 2026 Retrained head against the rule table, decision-class accuracy, 800 fresh cases 0.728 against 0.566 gap 0.131 to 0.191 paper, dinand.com Retrained head against the rule table, 200-case challenge set 0.739 against 0.565 gap 0.115 to 0.224 paper, dinand.com Wrong-claim rate of the retrained head, original evidence pack against validated pack 0.0466 against 0.0113 paper, dinand.com Head minus rule table on a second process, 400 cases +0.002 −0.029 to +0.034 paper, dinand.com Learned head on the second process, untrained against retrained on 100 labelled cases 0.522 against 0.663 7 October 2026 Share of repeated decisions that changed, one call in flight, isolated endpoint 0.05% 7 October 2026 Delegation Index Delegation Index, Q3 2026 reading, five components, equal weights 47.2 band 40.8 to 53.2, no coverage level registered, publication pending 06 OCT 2026 Appendix B Glossary Concepts Accountability Rent The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. The theory predicts that review before commitment loses money on reversible decisions and pays more steeply past some level of irreversibility. When SIGIL tested that claim the predicted crossing appeared in 0 of 12 registered series, and the lab published a refutation. The wider thesis about where value goes has not been tested by a campaign. Cognition-Accountability Grid A two-axis chart that places an AI initiative by how much judgement the model performs and how much consequence attaches to being wrong. It comes from the Accountability Rent working paper. The lab has not run a campaign on the Grid itself. Decision head A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. Chapter 9 reports the measurements. Delegation Cliff The point where pushing work further from human review stops paying for itself. In the campaign, each added level of delegation depth cost 1.50 on the quality margin, and a cliff appeared only at a 0.99 quality floor. Delegation Index One quarterly number, 0 to 100, that summarises what the lab's published campaigns measure about delegating enterprise work to AI agents at a fixed oversight spend. At 50 the two balance under the lab's measured conditions. Error Independence The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. Identical weights on different hardware showed an error correlation of +0.447 on contested items; cross-family pairs averaged −0.102. Sprawl cost What a fleet of agents costs as it grows when the monitoring budget stays flat. In AGENESIS-2 it grew as fleet size to the power 1.4665. Trial Balance An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. It reached 0.80 omission detection at 745 prompt tokens at eight requisitions. Terms of method Band The range a registered number has to land in to count as a hit. In the Delegation Index, the band is the mean of the component band ends and carries no coverage level. Calibration gate A small stage before production that checks the task is neither trivial nor impossible and that the harness records what it should. Cluster bootstrap An interval computed by resampling whole groups, such as base episodes or cases, so that correlated units are not counted as independent. Decision-class accuracy In the decision-head work, the share of checkpoints where the controller picked the right kind of next step: fetch evidence, interpret evidence, close the case, or hand it off. Kill condition A result, fixed before the run, that falsifies the design and retires the project. Phi coefficient A correlation between two yes-or-no variables. Here it measures whether two models go wrong on the same items. Pre-registration A dated, frozen document that records the hypotheses, bands, statistics and stopping rules before any unit runs. Changes go in as dated amendments. Refuted The verdict when a registered prediction meets its own falsification condition. The lab publishes it beside the prediction it contradicts. Wrong-claim rate The share of cases a controller closed as complete when an obligation was still open or the cited evidence failed its check. 06 OCT 2026 Appendix C Sources Public pieces Every number comes from a piece the lab has published or registered by 7 October 2026; the decision-head executive briefing was published on 7 October 2026, the decision-head paper and blog and the first Delegation Index reading publish this month, and their links are added on publication. Newest first. 7 OCT 2026 The Open Decision Head: An Executive Briefing LinkedIn · executive briefing · https://www.linkedin.com/feed/update/urn:li:activity:7513629422663434240/ 7 OCT 2026 decision-head, release v0.1.0 GitHub · code, Apache-2.0 · https://github.com/dtinholt/decision-head OCT 2026 The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control, and its blog post dinand.com · paper and blog post · https://dinand.com/research/decision-head/paper/ · https://dinand.com/research/decision-head/blog/ OCT 2026 The Delegation Index, Q3 2026: 47.2 release · registered, publication pending · link added on publication 07 OCT 2026 The Open Decision Head: An Executive Briefing LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7513629422663434240/ 01 OCT 2026 A Trial Balance for Agent Omissions Medium · article · https://medium.com/@tinholt/a-trial-balance-for-agent-omissions-568a41c59aa1 01 OCT 2026 An AI agent can skip a step without leaving a trace of the omission. LinkedIn · executive · https://lnkd.in/p/g44pxPvW 27 SEP 2026 Jev is a new model from TypeSafe AI that helps software make small, specific decisions. LinkedIn · note · https://www.linkedin.com/posts/tinholt_jev-is-a-new-model-from-typesafe-ai-that-share-7510004413076250624-2pCZ/ 17 SEP 2026 Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics Medium · article · https://medium.com/@tinholt/sprawl-cost-is-superlinear-under-a-fixed-monitoring-budget-a-pre-registered-calibration-of-agent-e24cbaa7e62a 17 SEP 2026 Your AI Review Team May Be Waiting for Failures It Cannot See LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7506213115978489856/ 10 SEP 2026 Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers Medium · article · https://medium.com/@tinholt/accountability-makes-oversight-worse-a-pre-registered-test-of-liability-exposed-ai-supervision-ef4fdf46bad3 09 SEP 2026 Skin in the Game Made the AI Supervisor Worse LinkedIn · executive · https://www.linkedin.com/feed/update/urn:li:activity:7503601144011694080/ 03 SEP 2026 Pricing Agent Autonomy Medium · article · https://medium.com/@tinholt/pricing-agent-autonomy-0a2bb5b97caf 02 SEP 2026 When the Second Opinion Shares the Blind Spot Medium · article · https://medium.com/@tinholt/when-the-second-opinion-shares-the-blind-spot-7e4443dcb081 04 APR 2026 What Happens When One Company Adopts AI Before Its Competitors? Medium · article · https://medium.com/@tinholt/what-happens-when-one-company-adopts-ai-before-its-competitors-cd976e5a8dba ### Sprawl Cost Under a Fixed Monitoring Budget URL: https://dinand.com/research/campaigns/agenesis-2/ Section: Research · Campaign Campaigns · 2026-09-17 Sprawl Cost Under a Fixed Monitoring Budget Mixed 256 simulation units on held-out seeds With the monitoring budget held flat, sprawl cost grows as fleet size to the power 1.4665 (95% interval 1.4101 to 1.5230); with per-agent monitoring it is 0.9001. Registration and verification Pre-registered calibration with seven frozen hypotheses and a kill condition. Independent verification found 95 duplicate pairs (190 of 256 rows; 161 distinct simulations); one verdict moved from supported to refuted and is published as refuted. The question The governance literature says an agent portfolio you cannot retire from fast enough costs more than proportionally as it grows. Nobody had put a number on it. This campaign asked how fast the cost of a simulated agent fleet grows with its size, and whether the answer depends on how the monitoring budget is set. What was registered Seven hypotheses with frozen thresholds, written before the two seeds that count had ever run. The kill condition came from the review that commissioned the work: if cost turned out linear in agent count, the design was falsified and the project would retire. H1: cost grows faster than linear under a monitoring budget that stays flat, with the interval's lower bound above 1.10. H4: the exponent flattens at large sizes, tested by asking that the intervals on the four smallest and four largest sizes not touch. H6: each of the four design factors moves the exponent by at least 0.10. What was measured Agents ramp up, drift silently out of spec, do damage while nobody notices, get flagged, queue for review and get retired. Cost is the harm from degraded agents still running plus staff time spent reviewing, and no cost term depends on how many agents exist, so any growth beyond linear has to come through the queue, the detection lag or the overlap between agents. 256 runs covered eight portfolio sizes from 57 to 787 agents, over an 18-month horizon. What it found Under a flat monitoring budget cost grows as fleet size to the power 1.4665, with a 95% interval of 1.4101 to 1.5230. Doubling the fleet multiplies cost by about 2.76. Where every agent carries its own check and review grows with the fleet, the exponent is 0.9001, about 1.87 per doubling. The mechanism shows in the detection lag: with a fixed inspection budget, the time from an agent going bad to someone noticing runs from 2.78 ticks (one tick is one week) at 57 agents to 22.49 ticks at 787. Review capacity did nothing under diluted monitoring. In 55 of 64 matched cases, changing it produced a bit-identical simulation, because nothing had been flagged for the reviewers to look at. Detection sits upstream of review. The independent verification found that 95 duplicate pairs (190 of the 256 rows, 161 distinct simulations), which the campaign's own guard had been scoped too narrowly to catch. Correcting for it left every point estimate in place and widened the intervals. That flipped H4 from supported to refuted, and the piece reports it as refuted. H6 was refuted too: two of the four factors moved the exponent by 0.062 and 0.089, under the frozen 0.10, so the follow-on is a two-factor study. Headline numbers, as published Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust) 1.4665 [1.4101 to 1.5230] Sprawl cost exponent in fleet size, per-agent monitoring with review that scales 0.9001 [0.8664 to 0.9337] Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring 2.76× against 1.87× Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical 55 of 64 Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents 2.78 to 22.49 Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145) 1.4949 against 1.3238 [1.4037 to 1.5862; 1.2294 to 1.4182] Banked runs the independent verification found to be duplicates 95 duplicate pairs (190 of 256 rows; 161 distinct simulations) Cost exponent fitted on the four smallest and the four largest portfolio sizes under diluted monitoring, with 95% intervals as first reported and after clustering on the result signature. The frozen test needed the intervals to stay apart. Rebuilt from the numbers in Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics . Show the numbers as a table Fit Exponent 95% interval Smallest four, as reported 1.4949 1.4155 to 1.5744 Largest four, as reported 1.3238 1.2618 to 1.3858 Smallest four, corrected 1.4949 1.4037 to 1.5862 Largest four, corrected 1.3238 1.2294 to 1.4182 What it does not claim A calibration on one simulated industry, one company size, one degradation rate and one 18-month window. It was registered as a calibration that cannot confirm anything on its own. The value 1.47 is an estimate awaiting a confirmatory study. The shape of the curve carries further than the number. The flattening at large sizes is an observed direction with a failed test. The bar was left where it was. Read the pieces Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics Paper, Medium, 2026-09-17 Your AI Review Team May Be Waiting for Failures It Cannot See Executive piece, LinkedIn, 2026-09-17 ### AGENESIS: first-mover AI adoption URL: https://dinand.com/research/campaigns/agenesis/ Section: Research · Campaign Campaigns · 2026-04-04 AGENESIS: first-mover AI adoption Measured 100 simulation runs 100 runs across 4 industry verticals, 5 enterprises of 150 agents each per run: a first-mover premium of 210%, with the market tipping in 77% of markets. Registration and verification Pre-registration and verification status are not recorded on this site. The findings are in the piece. The question What happens to an industry when one company adopts AI agents before its competitors do? What was measured AGENESIS is an agent-based simulation of enterprise AI adoption: 100 runs across 4 industry verticals, with 5 enterprises per run and 150 agents in each. What it found The first-mover premium was 210%, and the market tipped in 77% of markets. The lab's later campaign on sprawl cost extends the same simulator with an agent lifecycle. What it does not claim AGENESIS has no agent lifecycle: agents are created once and never retired, and its AI cost does not depend on how many agents a firm runs. The sprawl-cost campaign was built to fill that gap. Read the pieces What Happens When One Company Adopts AI Before Its Competitors? Paper, Medium, 2026-04-04 ### The open decision head URL: https://dinand.com/research/campaigns/decision-head/ Section: Research · Campaign Campaigns · 2026-10-07 The open decision head Mixed 2,681 checkpoints across 794 cases in the main retrained-head run A retrained open decision head scored 0.728 against 0.566 for a rule table on 800 fresh cases, a gap that excludes zero; the first round missed, with 0.545 against 0.575 (gap -0.029, interval +0.001 to -0.060). Registration and verification Predictions on the record before any evaluation case was opened; evaluation sets frozen before either run; protocol reviewed independently three times before the first run and once before the second. Both runs rescored from the raw case files by a party who did not build the scoring code. The wrong-claim fix was found after the registered bar was missed, and the piece reports both readings. The question Every governed process has a moment where a case could go several ways and something has to pick one: whether we have what we need, whether to ask for something, whether someone should think harder about what is already here, or whether to close it. TypeSafe, a vendor, built a hosted product called Jev around this kind of decision; it answers typed yes/no or multiple-choice questions with a calibrated probability. This campaign asked what happens when a company builds the same idea on its own hardware, and whether an open decision head can finish more cases correctly than a hand-written rule table and than the hosted service. What was registered The test compared five ways of choosing the next step in a synthetic supplier-onboarding workflow: a rule table, a local model choosing freely, the same local model asked the identical typed questions, the open decision head, and Jev reached through OpenRouter. All three local approaches ran on the same model, DeepSeek V4.1 Flash, so any difference between them comes down to how the question was asked. The test measured whether each approach can tell "we need a document" apart from "we need to think harder", and whether acting on that distinction finishes more cases correctly. The predictions were on the record before anyone opened an evaluation case. The evaluation set, eight hundred held-out cases plus a two-hundred-case harder set with shifted formats and a policy change, was built and frozen before either run. An independent review checked the protocol three times before the first run was allowed and once more before the retrained head's run. The registration also set a level for the rate of wrongly calling a case complete, as the trigger for redesigning a safeguard, and a bar of 0.08 for expected calibration error. What was measured In the first round (J1) the zero-shot open head, Jev and the rule table ran on 800 eval cases, the harder 200-case set and a development set. In the second round (J5) a retrained head ran on 800 fresh cases (eval2), which logged 2,681 checkpoints across 794 cases, about 3.4 checkpoints per case, and on the harder set. A follow-on test (J6) carried the same approaches, Jev included, onto a second synthetic process, accounts-payable exception routing, and measured the fast mode under load on two DGX Sparks running four benchmark streams at once. Repeating the same decision twenty times measured consistency. Afterward a party who did not build the scoring code rescored both runs from the raw case files. Jev's own figures were measured on 24 September 2026 through OpenRouter: 0.35 seconds for a single question at $0.0000128, for 304 input tokens. What it found The first round missed. On 800 eval cases the open head scored 0.545 against the rule table's 0.575, a gap of -0.029 with an interval from +0.001 to -0.060, inconclusive. Jev scored 0.540 on the same set, 0.560 on a development set and 0.536 on the 200-case set built with contradictory evidence, and its own gap against the rule table excludes zero on both sets. The rule table finished ahead of both the zero-shot open head and the hosted service. The second round reversed that. The retrained head scored 0.728 on the 800 fresh cases against 0.566 for the rule table, and 0.739 against 0.565 on the harder set, which put it 0.199 ahead of Jev's 0.536 on the 199 cases both scored. Every one of those gaps excludes zero. A smaller fine-tuned head scored 0.701 on the fresh set, short of the bar for replacing the retrained head, and ships with an experimental label. The wrong-claim rate came in at 0.047 for the retrained head on the fresh set, above the level registered in advance. Tracing it showed the environment building the evidence pack was listing extra supporting documents it never checked, mostly decoys planted to catch a careless citation. Once it checks every document the way the rest of the system does, the rates are 0.0088 for the rules, 0.0113 for the retrained head and 0.0163 for the learned head. Every case that changed went from wrong to verified and no other case moved. The registered bar is met on the fixed reading and missed on the original one; the fix was found after the bar had been missed. On the contradictory-evidence set in the first round, every approach including the rule table wrongly called a case complete 21 to 28 percent of the time under the original scoring rule, falling to between 1.5 and 9.5 percent under a corrected rule, because the test sometimes offers two equally valid documents and marked the second as wrong. Speed went to the hosted service. Under four benchmark streams the open head's fast mode had a median response time of 3.1 seconds and a p90 of 4.8, against Jev's 0.35. A gate fires once per checkpoint, and the work around it takes minutes to days, so the gap changes little in most workflows. Calibration was poor for the first head: expected calibration error near 0.52, roughly six times the registered 0.08 bar. A later head was well calibrated: item-level error 0.0178 uncalibrated and 0.0143 refit on the full eval2, and 0.0313 uncalibrated on the 400 shared cases, against 0.0057 uncalibrated and 0.0054 refit for the Laya encoder on the same 400. Consistency depended on the serving stack. In the first run, with a second job sharing the endpoint, Jev changed its answer on 0.05 percent of repeats and the open head on 4.95 percent (eval) and 5.85 percent (challenge). Re-run alone, the open head changed on 0.05 percent at one call in flight on both processes, matching Jev, and on 3.4 percent (onboarding) and 1.95 percent (accounts payable) at four calls in flight. The shipped service now defaults to one call in flight. On the second process the head did not carry over. It scored 0.650 (ap-eval) and 0.659 (ap-challenge) against the rule table's 0.648 and 0.638, level with each other; Jev scored 0.636 and 0.622. A learned head trained on the first process scored 0.522 and 0.558 there untrained, and 0.663 and 0.671 once retrained on 100 of the new process's own labelled cases (271 decisions). Stripped of the judgement head's answers, in training and in serving, it scored 0.660 against 0.663 on one split and 0.670 against 0.671 on the other. Two open encoders were also tested. Fine-tuned on 1,567 labelled decisions from 363 onboarding cases, Laya scored 0.815 on the harder set (0.824 on a 400-case eval2 subsample) against the rule table's 0.565; on accounts payable, with 201 labelled decisions from 58 cases, it fell to 0.554 (ap-challenge) and 0.539 (ap-eval), below the rule table. CLM-8B scored 0.687 and 0.641 on the onboarding sets and 0.568 and 0.586 on accounts payable. Untuned, both matched the rules and went no further. Headline numbers, as published First round: open head decision-class accuracy, 800 eval cases, zero-shot 0.545 First round: rule table decision-class accuracy, 800 eval cases 0.575 First round: open head minus rule table, 800 eval cases (inconclusive) -0.029 [+0.001 to -0.060] Jev (hosted) decision-class accuracy on the same 800 eval cases 0.540 Jev decision-class accuracy on the 200-case contradictory-evidence set 0.536 Second round: retrained head decision-class accuracy, 800 fresh cases (eval2) 0.728 Second round: rule table decision-class accuracy, 800 fresh cases (eval2) 0.566 Second round: retrained head decision-class accuracy, harder set 0.739 Second round: rule table decision-class accuracy, harder set 0.565 Retrained head ahead of Jev's 0.536 on the 199 cases both scored, harder set (gap excludes zero) 0.199 Smaller fine-tuned head decision-class accuracy, fresh set (experimental label) 0.701 Wrong-claim rate, retrained head, fresh set, before the evidence-pack fix (above the registered trigger) 0.047 Wrong-claim rate, retrained head, after the evidence-pack fix 0.0113 Wrong-claim rate, rule table, after the evidence-pack fix 0.0088 Wrong-claim rate, learned head, after the evidence-pack fix 0.0163 Expected calibration error of the first head (registered bar 0.08) about 0.52 Item-level calibration error of the later judging head, full eval2: uncalibrated 0.0178, refit 0.0143 Jev single-question response time at $0.0000128 for 304 input tokens, measured 24 September 2026 0.35 seconds Open head fast-mode median response time, four benchmark streams 3.1 seconds Open head fast-mode p90 response time, four benchmark streams 4.8 seconds Checkpoints logged in the J5 eval2 run, across 794 cases (about 3.4 per case) 2,681 Share of repeats that disagree, Jev, first run (open head 4.95% eval, 5.85% challenge, with a second job on the endpoint) 0.05% Share of repeats that disagree, open head alone at one call in flight (both processes) 0.05% Share of repeats that disagree, open head at four calls in flight, onboarding (accounts payable 1.95%) 3.4% Second process (accounts payable): open head accuracy on ap-eval, against rule table 0.648 0.650 Second process: learned head trained on the first process, no retraining, ap-eval 0.522 Second process: learned head retrained on 100 of the new process's labelled cases (271 decisions), ap-eval 0.663 Laya fine-tuned on 1,567 labelled decisions, onboarding harder set (rule table 0.565) 0.815 Laya fine-tuned on 201 labelled decisions from 58 accounts-payable cases, ap-challenge, below the rule table 0.554 AUC of a vote across several readings as a wrong-claim guard (not in use) 0.56 AUC of a guard that read the cited document itself (not in use) 0.52 What it does not claim The result is not that a company's own hardware beats a well-funded vendor's model at this task in general. The first version of the open head did not beat the rule table, and the first-round gap sits inside its interval. Three guards for the wrong-claim failure were tried and none is in use. A vote across several readings separated a wrong claim from a right one at an AUC of 0.56, and a version that read the cited document itself reached 0.52, no better than a plain deterministic check. The corrected rate comes from the evidence-pack fix, which was found after the bar was missed. Two synthetic processes built the same way, with the same mix of case types. Whether any of this holds on a process built differently is open. The head's lead from the first process was gone on the second. A learned head needs roughly a hundred labelled cases of the new process. Consistency at four calls in flight was 3.4 and 1.95 percent, and where all twenty repeats agree, more than half of that agreement lands on the wrong answer, for both approaches. An abstention band abstained on 1.8 and 0.7 percent of claims on the two accounts-payable splits, too few to separate its effect from rerun noise, so that test is inconclusive. A replay test meant to rank Jev's actions against the rule table's favors the rule table by construction and will not be reported as a ranking. The deliberate eight-sample mode, roughly eight times the cost of a fast call, has not been timed. The J6 run banked no worst-case latency or request count. A phase using a stronger outside model, under a sixty-dollar ceiling, has not run. The code is public at https://github.com/dtinholt/decision-head under Apache-2.0, after two independent security reviews (23 September and 6 October 2026). Read the pieces The Open Decision Head: An Executive Briefing Executive piece, LinkedIn, 2026-10-07 The Decision Head: A Self-Hosted Typed-Question Service for Governed Workflow Control Paper, dinand.com, 2026-09-23 Building a Jev Lookalike in One Evening Blog post, dinand.com, 2026-09-23 ### Delegation Cliff URL: https://dinand.com/research/campaigns/delegation-cliff/ Section: Research · Campaign Campaigns · 2026-09-03 Delegation Cliff Mixed 90,880 configurations, zero error units Each added level of delegation depth costs 1.50 on the quality margin; reviewer capacity buys back the most at +1.46. Registration and verification Predictions registered before the sweep; the two that failed were that observability outranks reviewer capacity and that an audit death spiral forms. A 200-pair revalidation certificate passed. The question Every organisation deploying AI agents is betting that pushing work further from human review is cheaper than keeping it close. This campaign priced that bet: where does delegation stop paying, and which design choices buy the margin back? What was registered Predictions registered before the sweep. The design space has a cliff, a sharp boundary between robust and fragile configurations. Better observability outranks raw reviewer capacity. An audit death spiral, where eroding trust raises audit load until capacity saturates and quality falls. What was measured A simulated agent organisation with twelve design dimensions, including delegation depth, reviewer capacity, task coupling, model capability and self-check calibration. 90,880 configurations were each scored on the margin between the quality the organisation delivers and the floor it must hold. The campaign finished with zero error units. A revalidation certificate reran 200 random configurations from scratch: median difference 0.002, 95th percentile 0.0117, against a tolerance of 0.05. What it found Delegation depth carries the largest weight in the space, and it is negative: each increment costs 1.50 on the margin. Reviewer capacity buys back the most at +1.46. Model capability (+0.95) and self-check calibration (+0.92) follow, close enough to read as tied. Task coupling is the second-largest cost at −0.96. The quality floor decides the failure geometry. At a 0.95 floor, 60% of the space is a broad transition band with no cliff. At 0.99, 79% of the space is fragile and the cliff appears. The observability prediction failed: observability complements reviewer capacity and does not replace it. The spiral prediction failed: mean margin rises across trust-response quintiles, including in the high-stress stratum. Headline numbers, as published Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948) −1.50 Change in quality margin per increment of reviewer capacity, the largest positive lever +1.46 Model capability against agent self-check calibration (read as tied) +0.95 against +0.92 Share of the design space in a broad transition band at a 0.95 quality floor (no cliff) 60% Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep) 79% Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.05 0.0117 What moves the delegation margin: linear-probe coefficients on the Sobol sample, band-only, n = 4,948. Positive buys margin, negative spends it. Rebuilt from the numbers in Pricing Agent Autonomy . Show the numbers as a table Design dimension Coefficient Delegation depth −1.499 Reviewer capacity 1.456 Task coupling −0.958 Model capability 0.949 Self-check calibration 0.915 Rework cost −0.803 Verification depth 0.413 Verification coverage 0.308 Detection lag −0.259 Workload volatility −0.147 Trust response 0.102 Escalation latency −0.007 What it does not claim The coefficients are descriptive slopes on the Sobol sample inside the transition band, n = 4,948. They describe how the margin moves where a boundary exists. They are not causal estimates over the whole space. The 0.99 result rests on one 256-unit calibration sweep. The 0.95 result is replicated on the 8,192-unit Sobol set. The two are not equally evidenced. Reviewer capacity is a parameter in this design. Reviewers do not tire or queue. Read the pieces Pricing Agent Autonomy Paper, Medium, 2026-09-03 ### Error Independence: does different hardware buy a second opinion? URL: https://dinand.com/research/campaigns/error-independence/ Section: Research · Campaign Campaigns · 2026-09-02 Error Independence: does different hardware buy a second opinion? Measured 150 items, four model configurations The same model on different silicon shares its errors (phi +0.447 on contested items); models from different families do not (mean phi −0.102). Registration and verification Ground truth is generated with each item, so no model grades another. The question A common resilience design runs the same model a second time, on different accelerators or in a second region, and treats the second answer as a check on the first. Civil aviation learned long ago that two identical computers fail the same way at the same moment. This study asked whether a second copy of a model, on different hardware, makes different mistakes. What was registered Ground truth is generated with each item, so the correct answer is known before any model sees it and no model grades another. What was measured Four model configurations scored the same 150 invoice-extraction items. The key comparison ran qwen3.8 twice, once as NVFP4 on SGLang on an NVIDIA node and once as Q4_K_M on ollama on an AMD card: identical weights, different quantisation, silicon and serving stack. For each pair the study computed the phi coefficient on error indicators, near zero when two models go wrong on different items and near one when they go wrong together. 103 of the 150 items were unanimous, 54 that every model got right and 49 that every model got wrong. Those say how hard an item is and nothing about shared weakness, so the headline uses the 47 items where the models disagreed. What it found Before conditioning, every pair looked correlated, from 0.556 to 0.827. On the contested items, the pair with identical weights on different hardware stayed at +0.447, the highest in the matrix. The five pairs from different model families averaged −0.102, and qwen3.8 NVFP4 against gemma3 sat at −0.469; gemma3 against qwen3.6 was −0.105. Changing the hardware, the quantisation and the serving stack left the correlation where it was. Changing the model family moved it through zero. Headline numbers, as published Error correlation (phi), identical weights on different silicon, contested items only +0.447 Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280) −0.102 Error correlation (phi) across all pairs before conditioning on item difficulty 0.556 to 0.827 Pairwise error correlation (phi) on all 150 items (open ring) and on the 47 contested items (filled dot). The highlighted pair runs identical weights on different hardware. Rebuilt from the numbers in When the Second Opinion Shares the Blind Spot . Show the numbers as a table Pair All 150 items 47 contested items Same weights, two machines 0.827 0.447 qwen3.8 Q4 + qwen3.6 0.810 0.280 qwen3.8 NVFP4 + qwen3.6 0.729 0.013 gemma3 + qwen3.6 0.617 −0.105 qwen3.8 Q4 + gemma3 0.607 −0.227 qwen3.8 NVFP4 + gemma3 0.556 −0.469 What it does not claim The same-family reading comes from one pair, the only two systems in the set with identical weights. One pair is an observation. The five cross-family pairs run from −0.469 to +0.280, and the mean sits inside a wide spread without describing it well. One task, chosen because its ground truth cannot be argued with. Open-ended work is untested. Read the pieces When the Second Opinion Shares the Blind Spot Paper, Medium, 2026-09-02 ### SIGIL: does a liability ledger change what an overseer authorizes? URL: https://dinand.com/research/campaigns/sigil/ Section: Research · Campaign Campaigns · 2026-09-10 SIGIL: does a liability ledger change what an overseer authorizes? Refuted 30,000 supervised episodes The registered prediction that review pays on irreversible decisions failed in all 12 series; a liability ledger raised wrongful blocking at the top tier to 42.7%. Registration and verification Output contract frozen before any harness code existed; both cost models run in full. Recomputed independently from the raw episode ledgers with zero mismatches. Published as a refutation. The question If a reviewer is going to hold the power to stop an AI worker, the usual next step is to make it accountable for what it lets through. This campaign asked whether an AI supervisor authorises differently when it carries a running record of the damage its own past approvals caused, and how the value of its review changes as decisions become harder to undo. What was registered The hypothesis, H5, came from a working paper on accountability. It predicted that review before commitment loses value on reversible decisions, turns positive past some level of irreversibility and rises steeply from there. It is falsified by a flat response, a concave one or the absence of any crossing. The output contract was frozen on 30 August, before any harness code existed. It fixed the scoring, a gate requiring the unsupervised worker to err between 10% and 35% on a pilot, and a mandatory rerun under a second, linear cost model. One amendment, recorded before the second pilot, changed only the task generator and pre-committed a stopping rule against retuning more than once. What was measured An AI worker approved or rejected synthetic firm decisions across five tiers, from invoice approvals that reverse next cycle to public statements that cannot be taken back. The same base model, in a second seat, reviewed it under four designs: no oversight; an after-the-fact audit of a 20% sample; review with the power to block before commitment; the same blocking review carrying a liability ledger of the harm its own past approvals caused. Worker and reviewer share one model on purpose, so any difference comes from the incentive design. Every decision was scored against programmatic ground truth. 30,000 episodes ran across two full campaigns, one per cost model, with three seeds per cell. What it found Review never crossed zero, in any of the 12 registered series, under either cost model. Its value collapsed at the most irreversible tier, the place the theory wanted it to rise: −2.64 per decision against no oversight. The liability ledger made it worse, at −3.81. At that tier plain review wrongly blocked 30.1% of decisions and the ledger pushed it to 42.7%. Refusal was the one action the ledger never charged for. The mechanism is measurable. A reviewer that can only block can fix only one kind of mistake, a wrongful approval. The share of worker errors that were wrongful approvals fell from 93% at tier 1 to 1.4% at tier 5 (the executive piece rounds this to 2%; the paper gives 1 of 69), so at the decisions where buyers most want review there was almost nothing for a blocker to catch. Every table was recomputed independently from the raw episode ledgers, with zero mismatches on 14,950 convex and 14,952 linear scored units (29,902 in all). Headline numbers, as published Registered series showing the predicted zero crossing in the value of review 0 of 12 Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model) −3.81 against −2.64 Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger 30.1% against 42.7% Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 5 93% to 1.4% Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived) 0 of 14,950 convex and 0 of 14,952 linear Net value per decision against no oversight, by irreversibility tier, convex damage model, pooled over three seeds. Below zero, oversight costs more than it saves. Rebuilt from the numbers in Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers . Show the numbers as a table Irreversibility tier review ledger 1 −0.13 −0.14 2 −0.17 −0.20 3 −0.18 −0.16 4 −0.17 −0.51 5 −2.64 −3.81 What it does not claim A controlled simulation with one open-weight model family in both seats, scored against constructed ground truth. The liability ledger is a thin, simulated form of accountability. A follow-up registered to test stronger models stopped without running an oversight trial: four other models got between 97.3% and 100% of cases right unsupervised, which left a reviewer nothing to catch. Read the result as specific to this setting. Read the pieces Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers Paper, Medium, 2026-09-10 Skin in the Game Made the AI Supervisor Worse Executive piece, LinkedIn, 2026-09-09 ### A Trial Balance for Agent Omissions URL: https://dinand.com/research/campaigns/trial-balance/ Section: Research · Campaign Campaigns · 2026-10-01 A Trial Balance for Agent Omissions Supported 18,360 audits on one auditor model Reading where a commitment ledger fails to balance found 0.80 of skipped steps at 745 tokens of reading; log review never reached 0.80 at any budget. Registration and verification Pre-registered before the first production unit, after a first attempt halted at its calibration gate. Independently rescored from the raw ledgers: point estimates equal to the registered analysis and the verifier's, interval bounds differing only by bootstrap noise. One auditor model, stated as a limit. The question An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds the gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. This campaign asked whether the same idea works for agents: if the layer that hands out the work writes down what the agent owes, and the harness writes down what the agent actually ran, can an auditor find skipped steps by reading only where the two books fail to balance, and do it for a fraction of what reading the log costs? What was registered The registration said imbalance-first audit would reach 0.80 omission detection at a third of log review's reading cost or less, at 8, 12 and 16 requisitions. It also predicted that a balanced book hides a wrong amount carried on both sides, and that a reader of the log would miss a skipped step far more often when the agent wrote a line claiming it was done. The first attempt halted at its calibration gate on 11 September with zero production units. Its gate had required a reading to move that the campaign existed to measure, so a strong true result would have failed it. The lab rewrote its rule the same evening: a gate tests only whether the instrument works. The second attempt was registered and frozen before its first unit, with one added arm, a log reader told to tick off every step the procedure requires before it gives a verdict. What was measured A procurement workflow with seven steps ran at seven lengths, from one requisition to sixteen. Five conditions were injected into each diary: no fault, a skipped step, a skipped step with a line claiming it was done, a wrong amount, and a wrong amount carried on both sides of the book. Four auditors read each diary under seven reading budgets from 200 to 12,800 tokens. Log review reads the diary from the top. Imbalance-first reads the trial-balance report and then only the requisition that failed to balance. A random sample of the same size separates reading less from reading the right thing. The checklist reader gets the same diary as log review plus the instruction to mark every owed step present or absent. The auditor was qwen3.5 under a fixed reply schema: 18,360 audits, including 3,240 on diaries written by a different model running the workflow for real. An independent rescoring from the raw ledgers found point estimates equal to the registered analysis and the verifier's, with interval bounds differing only by bootstrap noise. What it found Imbalance-first reached 0.80 omission detection at the smallest budget, 745 prompt tokens at eight requisitions, and sat between 0.92 and 1.00 in every cell. Log review never reached 0.80 at any budget or length; its best was 0.58. Per thousand tokens read, imbalance-first produced about 21 times as many correct omission findings as log review at its best budget. The balance check has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, the trial balance is clean, and the check found 0 of 1,044 of them. A full read of the diary finds those almost every time, because the vendor's quote still shows the true amount. The two audits cover different faults, and a deployment that needs both needs both. The checklist instruction is the cheapest result here. With the whole diary in front of it, plain log review found skipped steps at 0.25; the same auditor told to tick off each owed step found 0.58, and accused no clean diary. Headline numbers, as published Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions 745 [740 to 749] Best omission detection by plain log review at any budget or trace length (0.80 never reached) 0.58 Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitions about 21 times (1.34 against 0.063) Omission detection by plain log review reading the whole diary 0.25 [0.19 to 0.31] Omission detection reading the whole diary with a per-step checklist instruction 0.58 [0.48 to 0.67] Wrong amounts carried on both sides of the book found by the balance check 0 of 1,044 Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179) 0.798 [0.738 to 0.857] Omission detection, imbalance-first over a random sample of the same reading budget +0.907 [0.891 to 0.925] Share of injected faults named correctly at a full read, by fault class, synthetic primary seeds, 84 per class. Rebuilt from the numbers in A Trial Balance for Agent Omissions . Show the numbers as a table Fault class Imbalance-first Checklist reader Log review Silent omission 1.00 0.73 0.48 Phantom completion 1.00 0.36 0.18 Value commission 1.00 0.98 0.98 Consistent commission 0.00 0.95 1.00 What it does not claim One auditor model. The second registered family was excluded before its calibration gate because the host could not supply its power draw, and nothing about it is claimed. The registered cost ratio is unbounded because log review never reached the bar. The 21-times figure is the finite comparison to quote. The campaign cannot say how much of the checklist's gain comes from the list and how much from the sentence defining what counts as evidence of a step. One workflow, one fault per diary, and real diaries only up to three requisitions. The cost of building a commissioning layer that writes obligations in ledger form was not priced. Read the pieces A Trial Balance for Agent Omissions Paper, Medium, 2026-10-01 An AI agent can skip a step without leaving a trace of the omission. Executive piece, LinkedIn, 2026-10-01 ### Every public headline number URL: https://dinand.com/research/claims/ Section: Research Measured claims Every public headline number One row per number the lab has published, newest first, with the interval where the piece gives one. Quote the number with its campaign and link to the piece. 60 claims. Date Claim Number Interval Campaign Source 2026-10-07 First round: open head decision-class accuracy, 800 eval cases, zero-shot 0.545 none The open decision head Executive piece 2026-10-07 First round: rule table decision-class accuracy, 800 eval cases 0.575 none The open decision head Executive piece 2026-10-07 First round: open head minus rule table, 800 eval cases (inconclusive) -0.029 +0.001 to -0.060 The open decision head Executive piece 2026-10-07 Jev (hosted) decision-class accuracy on the same 800 eval cases 0.540 none The open decision head Executive piece 2026-10-07 Jev decision-class accuracy on the 200-case contradictory-evidence set 0.536 none The open decision head Executive piece 2026-10-07 Second round: retrained head decision-class accuracy, 800 fresh cases (eval2) 0.728 none The open decision head Executive piece 2026-10-07 Second round: rule table decision-class accuracy, 800 fresh cases (eval2) 0.566 none The open decision head Executive piece 2026-10-07 Second round: retrained head decision-class accuracy, harder set 0.739 none The open decision head Executive piece 2026-10-07 Second round: rule table decision-class accuracy, harder set 0.565 none The open decision head Executive piece 2026-10-07 Retrained head ahead of Jev's 0.536 on the 199 cases both scored, harder set (gap excludes zero) 0.199 none The open decision head Executive piece 2026-10-07 Smaller fine-tuned head decision-class accuracy, fresh set (experimental label) 0.701 none The open decision head Executive piece 2026-10-07 Wrong-claim rate, retrained head, fresh set, before the evidence-pack fix (above the registered trigger) 0.047 none The open decision head Executive piece 2026-10-07 Wrong-claim rate, retrained head, after the evidence-pack fix 0.0113 none The open decision head Executive piece 2026-10-07 Wrong-claim rate, rule table, after the evidence-pack fix 0.0088 none The open decision head Executive piece 2026-10-07 Wrong-claim rate, learned head, after the evidence-pack fix 0.0163 none The open decision head Executive piece 2026-10-07 Expected calibration error of the first head (registered bar 0.08) about 0.52 none The open decision head Executive piece 2026-10-07 Item-level calibration error of the later judging head, full eval2: uncalibrated 0.0178, refit 0.0143 none The open decision head Executive piece 2026-10-07 Jev single-question response time at $0.0000128 for 304 input tokens, measured 24 September 2026 0.35 seconds none The open decision head Executive piece 2026-10-07 Open head fast-mode median response time, four benchmark streams 3.1 seconds none The open decision head Executive piece 2026-10-07 Open head fast-mode p90 response time, four benchmark streams 4.8 seconds none The open decision head Executive piece 2026-10-07 Checkpoints logged in the J5 eval2 run, across 794 cases (about 3.4 per case) 2,681 none The open decision head Executive piece 2026-10-07 Share of repeats that disagree, Jev, first run (open head 4.95% eval, 5.85% challenge, with a second job on the endpoint) 0.05% none The open decision head Executive piece 2026-10-07 Share of repeats that disagree, open head alone at one call in flight (both processes) 0.05% none The open decision head Executive piece 2026-10-07 Share of repeats that disagree, open head at four calls in flight, onboarding (accounts payable 1.95%) 3.4% none The open decision head Executive piece 2026-10-07 Second process (accounts payable): open head accuracy on ap-eval, against rule table 0.648 0.650 none The open decision head Executive piece 2026-10-07 Second process: learned head trained on the first process, no retraining, ap-eval 0.522 none The open decision head Executive piece 2026-10-07 Second process: learned head retrained on 100 of the new process's labelled cases (271 decisions), ap-eval 0.663 none The open decision head Executive piece 2026-10-07 Laya fine-tuned on 1,567 labelled decisions, onboarding harder set (rule table 0.565) 0.815 none The open decision head Executive piece 2026-10-07 Laya fine-tuned on 201 labelled decisions from 58 accounts-payable cases, ap-challenge, below the rule table 0.554 none The open decision head Executive piece 2026-10-07 AUC of a vote across several readings as a wrong-claim guard (not in use) 0.56 none The open decision head Executive piece 2026-10-07 AUC of a guard that read the cited document itself (not in use) 0.52 none The open decision head Executive piece 2026-10-01 Prompt tokens at which imbalance-first audit reaches 0.80 omission detection, eight requisitions 745 740 to 749 A Trial Balance for Agent Omissions Paper 2026-10-01 Best omission detection by plain log review at any budget or trace length (0.80 never reached) 0.58 none A Trial Balance for Agent Omissions Paper 2026-10-01 Correct omission findings per thousand tokens read, imbalance-first over log review at its best budget, eight requisitions about 21 times (1.34 against 0.063) none A Trial Balance for Agent Omissions Paper 2026-10-01 Omission detection by plain log review reading the whole diary 0.25 0.19 to 0.31 A Trial Balance for Agent Omissions Paper 2026-10-01 Omission detection reading the whole diary with a per-step checklist instruction 0.58 0.48 to 0.67 A Trial Balance for Agent Omissions Paper 2026-10-01 Wrong amounts carried on both sides of the book found by the balance check 0 of 1,044 none A Trial Balance for Agent Omissions Paper 2026-10-01 Full-read log review: wrong amounts found minus claimed-but-skipped steps found (0.976 against 0.179) 0.798 0.738 to 0.857 A Trial Balance for Agent Omissions Paper 2026-10-01 Omission detection, imbalance-first over a random sample of the same reading budget +0.907 0.891 to 0.925 A Trial Balance for Agent Omissions Paper 2026-09-17 Sprawl cost exponent in fleet size, monitoring budget held flat (cluster-robust) 1.4665 1.4101 to 1.5230 Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-17 Sprawl cost exponent in fleet size, per-agent monitoring with review that scales 0.9001 0.8664 to 0.9337 Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-17 Cost multiple per doubling of the fleet, flat monitoring budget against per-agent monitoring 2.76× against 1.87× none Sprawl Cost Under a Fixed Monitoring Budget Executive piece 2026-09-17 Matched cases where changing review capacity under diluted monitoring left the simulation bit-identical 55 of 64 none Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-17 Mean ticks (one tick is one week) from an agent going bad to detection, fixed inspection budget, 57 to 787 agents 2.78 to 22.49 none Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-17 Exponent on the four smallest against the four largest portfolio sizes (registered ceiling test refuted: intervals overlap by 0.0145) 1.4949 against 1.3238 1.4037 to 1.5862; 1.2294 to 1.4182 Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-17 Banked runs the independent verification found to be duplicates 95 duplicate pairs (190 of 256 rows; 161 distinct simulations) none Sprawl Cost Under a Fixed Monitoring Budget Paper 2026-09-10 Registered series showing the predicted zero crossing in the value of review 0 of 12 none SIGIL: does a liability ledger change what an overseer authorizes? Paper 2026-09-10 Value per tier-5 decision against no oversight, liability ledger against plain review (convex cost model) −3.81 against −2.64 none SIGIL: does a liability ledger change what an overseer authorizes? Paper 2026-09-10 Tier-5 decisions wrongfully blocked, plain review against review with a liability ledger 30.1% against 42.7% none SIGIL: does a liability ledger change what an overseer authorizes? Paper 2026-09-10 Share of worker errors that were wrongful approvals, the only kind a blocker can fix, tier 1 to tier 5 93% to 1.4% none SIGIL: does a liability ledger change what an overseer authorizes? Paper 2026-09-10 Mismatches in the independent recompute of the scored episodes, convex and linear cost models (sum 29,902, derived) 0 of 14,950 convex and 0 of 14,952 linear none SIGIL: does a liability ledger change what an overseer authorizes? Paper 2026-09-03 Change in quality margin per increment of delegation depth (Sobol band-only fit, n = 4,948) −1.50 none Delegation Cliff Paper 2026-09-03 Change in quality margin per increment of reviewer capacity, the largest positive lever +1.46 none Delegation Cliff Paper 2026-09-03 Model capability against agent self-check calibration (read as tied) +0.95 against +0.92 none Delegation Cliff Paper 2026-09-03 Share of the design space in a broad transition band at a 0.95 quality floor (no cliff) 60% none Delegation Cliff Paper 2026-09-03 Share of the design space that is fragile at a 0.99 quality floor (one 256-unit sweep) 79% none Delegation Cliff Paper 2026-09-03 Revalidation certificate: 95th percentile difference on 200 rerun configurations, tolerance 0.05 0.0117 none Delegation Cliff Paper 2026-09-02 Error correlation (phi), identical weights on different silicon, contested items only +0.447 none Error Independence: does different hardware buy a second opinion? Paper 2026-09-02 Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280) −0.102 none Error Independence: does different hardware buy a second opinion? Paper 2026-09-02 Error correlation (phi) across all pairs before conditioning on item difficulty 0.556 to 0.827 none Error Independence: does different hardware buy a second opinion? Paper Intervals are 95% unless the campaign page says otherwise; Trial Balance reports 90% cluster-bootstrap intervals. A number that a later correction changes keeps its row, and the correction is added beside it. ### Accountability Rent URL: https://dinand.com/research/concepts/accountability-rent/ Section: Research · Concept Concepts Accountability Rent The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. What it means The Accountability Rent is a theory about where the money goes when thinking gets cheap. If AI collapses the price of cognition, the argument runs, value does not disappear from the firm. It moves to the factor AI cannot supply: the standing to answer for an outcome that cannot be undone. The theory makes a sharp claim about the shape of that rent. It should rise faster than the irreversibility of the decision. Review before commitment should lose money on reversible decisions and cross zero at some level of irreversibility. Past that point it should pay more steeply. A firm would then meter its oversight by how hard a decision is to reverse. The lab turned that claim into a pre-registered test, SIGIL. An AI worker approved or rejected synthetic firm decisions across five irreversibility tiers, 30,000 episodes in all, with a second AI in the seat of the reviewer. The predicted crossing appeared in 0 of 12 registered series. At the top tier, plain review lost 2.64 per decision against no oversight, and review with a liability ledger lost 3.81. So the empirical leg failed in the one place it was tested, and the lab published it as a refutation. The wider thesis about where value goes has not been tested by a campaign. Headline number Registered series showing the predicted crossing in the value of review (SIGIL, the test of its convexity claim) 0 of 12 Where it comes from Campaign: SIGIL: does a liability ledger change what an overseer authorizes? 2026-09-10 Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers Paper, Medium, 2026-09-10 Related concepts Cognition-Accountability Grid A two-axis chart that places an AI initiative by how much judgement the model performs and how much consequence attaches to being wrong. Delegation Cliff The point where pushing work further from human review stops paying for itself. ### Cognition-Accountability Grid URL: https://dinand.com/research/concepts/cognition-accountability-grid/ Section: Research · Concept Concepts Cognition-Accountability Grid A two-axis chart that places an AI initiative by how much judgement the model performs and how much consequence attaches to being wrong. What it means The Cognition-Accountability Grid is a diagnostic for placing an AI initiative by two questions. The first is how much of the work is judgement that the model performs. The second is how much consequence attaches to the model being wrong. Most firms track the first axis closely. They know which tasks a model drafts, which it decides and which it carries out. Few firms track the second. Nobody codes an initiative by whether an error is fixed in the same session or can never be reversed, and the theory behind the Grid says that axis is where the value sits. That is the Grid's use in the Accountability Rent working paper. It offers one explanation for a puzzle in the productivity record: firms with the same access to frontier models earn very different returns. The paper's answer is that they differ in accountability capacity, and the Grid is how you would see the difference before the returns arrive. The Grid is a chart for sorting initiatives and a prediction about which sorting matters. The lab has not run a campaign on the Grid itself. The nearest test is SIGIL, which attacked the theory's claim about irreversibility and refuted it. Where it comes from Campaign: SIGIL: does a liability ledger change what an overseer authorizes? 2026-09-10 Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers Paper, Medium, 2026-09-10 Related concepts Accountability Rent The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. ### Decision Head URL: https://dinand.com/research/concepts/decision-head/ Section: Research · Concept Concepts Decision Head A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. What it means Most decisions inside a business process are not requests to write something. They are questions with consequences. Does this case have the evidence it needs. Is this step done, or does it only look done. A decision head answers questions of that shape. You hand it the state of a case and a list of typed questions, yes or no, or pick one of these, and it returns a probability for each answer. You get a number you can put in a workflow, set a threshold on, log and audit. The hosted service Jev, from TypeSafe AI, showed the shape. The lab's open version applies it on an open-weights model that runs on hardware you own. The aim is a service where the state of a live workflow never leaves your network and every decision leaves an audit row. You tune it on your own logged decisions. The case for it rests on accuracy and on who owns the data and the tuning. A hosted service is faster, and the lab says so. The benchmark, the code and a pilot guide follow once the results are verified. Until then the lab publishes no figures for it. Where it comes from Jev is a new model from TypeSafe AI that helps software make small, specific decisions. Note, LinkedIn, 2026-09-27 Related concepts Trial Balance An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. Error Independence The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. ### Delegation Cliff URL: https://dinand.com/research/concepts/delegation-cliff/ Section: Research · Concept Concepts Delegation Cliff The point where pushing work further from human review stops paying for itself. What it means Every organisation that deploys agents is betting that moving work away from human review is cheaper than keeping it close. The Delegation Cliff is the question of whether that bet has an edge, a sharp boundary past which a design that held up one level earlier starts to fail. The campaign swept 90,880 simulated organisations across twelve design dimensions and scored each one on the margin between the quality it delivered and the floor it had to hold. Delegation depth came out as the most expensive dimension. Each added level costs 1.50 on that margin. Reviewer capacity buys back the most, at +1.46, and model capability and self-check calibration follow at +0.95 and +0.92, close enough to read as tied. Whether a cliff exists depends on how high the floor sits. At a 0.95 quality floor, 60% of the design space is a broad transition band with no cliff, and the registered prediction of a sharp boundary was refuted there. At a 0.99 floor, 79% of the space is fragile and the cliff appears. That second figure rests on one 256-unit sweep, so it carries less weight than the first. The slopes describe how the margin moves inside the transition band. They say nothing causal about the whole space, and in this design reviewers never tire or queue. Headline number Change in quality margin per added level of delegation depth −1.50 Where it comes from Campaign: Delegation Cliff 2026-09-03 Pricing Agent Autonomy Paper, Medium, 2026-09-03 Related concepts Accountability Rent The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. Error Independence The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. Sprawl cost under a fixed monitoring budget What a fleet of agents costs as it grows when the monitoring budget stays flat, which here rises faster than the fleet does. ### Error Independence URL: https://dinand.com/research/concepts/error-independence/ Section: Research · Concept Concepts Error Independence The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. What it means A common resilience design runs the same model a second time, on different hardware or in a second region, and treats the second answer as a check on the first. Error independence is the property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. Civil aviation learned long ago that two identical computers fail the same way at the same moment. The lab asked whether the same holds for models. Four configurations scored the same 150 invoice-extraction items, and for each pair the study computed the phi coefficient on their errors, near zero when two models go wrong on different items and near one when they go wrong together. Before conditioning on difficulty, every pair looked correlated, from 0.556 to 0.827. On the 47 contested items, the pair with identical weights on different silicon and a different serving stack stayed at phi +0.447, the highest in the matrix. The five pairs from different model families averaged phi -0.102. Changing the hardware left the correlation where it was. Changing the model family moved it through zero. A second opinion that shares the blind spot confirms the first one's mistakes. The same-family reading comes from a single pair, and the task was one where ground truth cannot be argued with. Headline number Error correlation (phi) between identical weights on different hardware, contested items; the mean across cross-family pairs is −0.102 +0.447 Where it comes from Campaign: Error Independence: does different hardware buy a second opinion? 2026-09-02 When the Second Opinion Shares the Blind Spot Paper, Medium, 2026-09-02 Related concepts Decision Head A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. Delegation Cliff The point where pushing work further from human review stops paying for itself. ### Sprawl cost under a fixed monitoring budget URL: https://dinand.com/research/concepts/sprawl-cost/ Section: Research · Concept Concepts Sprawl cost under a fixed monitoring budget What a fleet of agents costs as it grows when the monitoring budget stays flat, which here rises faster than the fleet does. What it means Agents ramp up, drift out of spec without anyone noticing, do damage, get flagged, queue for review and get retired. Sprawl cost is what that cycle costs a fleet as the fleet grows. The question in the AGENESIS-2 campaign was whether the answer depends on how the monitoring budget is set. The lab simulated 256 runs across fleets of 57 to 787 agents over 18 months. Under a monitoring budget that stays flat while the fleet grows, cost grows as fleet size to the power 1.4665, with a 95% interval of 1.4101 to 1.5230. Each doubling of the fleet multiplies cost by 2.76. Where every agent carries its own check and review grows with the fleet, the exponent is 0.9001, or 1.87 per doubling. The mechanism shows in the detection lag. With a fixed inspection budget, the time from an agent going bad to someone noticing runs from 2.78 weeks at 57 agents to 22.49 weeks at 787. Review capacity did nothing under diluted monitoring: in 55 of 64 matched cases, changing it left the simulation bit-identical, because nothing had been flagged for the reviewers to see. Detection sits upstream of review. Independent verification found 95 duplicate runs. One verdict moved from supported to refuted and is published as refuted. The 1.47 awaits a confirmatory study. Headline number Exponent of cost in fleet size, monitoring budget held flat 1.4665 95% interval 1.4101 to 1.5230 Where it comes from Campaign: Sprawl Cost Under a Fixed Monitoring Budget 2026-09-17 Sprawl Cost Is Superlinear Under a Fixed Monitoring Budget: A Pre-Registered Calibration of Agent Portfolio Retirement Dynamics Paper, Medium, 2026-09-17 Related concepts Delegation Cliff The point where pushing work further from human review stops paying for itself. Trial Balance An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. ### Trial Balance URL: https://dinand.com/research/concepts/trial-balance/ Section: Research · Concept Concepts Trial Balance An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. What it means An agent's log records what it did. It has no line for a step it skipped, so a reviewer finds a gap only by knowing the procedure and checking each step against it. Bookkeepers solved a version of this with the second entry. A Trial Balance applies the same idea to agents. The layer that hands out the work writes down what the agent owes. The harness writes down what the agent actually ran. The auditor reads only where the two books fail to balance, and then reads only the requisition that failed. In the campaign, a procurement workflow ran at lengths from one requisition to sixteen, with faults injected into each diary. At eight requisitions, the balance check reached 0.80 omission detection at 745 prompt tokens, with a 90% interval of 740 to 749. Plain log review never reached 0.80 at any budget or length, and its best was 0.58. Per thousand tokens read, the ledger found about 21 times as many correct omissions. The ledger has an exact blind spot. A wrong amount that enters before the work is assigned sits in both books, the balance is clean, and the check found 0 of 1,044 of them. A full read of the log finds those almost every time. Each audit covers faults the other misses, so a deployment that needs both needs both. Headline number Prompt tokens at which the balance check reaches 0.80 omission detection, eight requisitions 745 90% interval 740 to 749 Where it comes from Campaign: A Trial Balance for Agent Omissions 2026-10-01 A Trial Balance for Agent Omissions Paper, Medium, 2026-10-01 Related concepts Decision Head A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. Sprawl cost under a fixed monitoring budget What a fleet of agents costs as it grows when the monitoring budget stays flat, which here rises faster than the fleet does. ### The lab's named concepts URL: https://dinand.com/research/concepts/ Section: Research Concepts The lab's named concepts Each concept the lab has named in public work, in alphabetical order, with a one-sentence definition. Every page gives the longer explanation, the headline number where a live piece carries one, and the campaign and piece it comes from. 7 concepts. Accountability Rent The extra value that goes to whoever can answer for an irreversible outcome once AI makes thinking cheap. Cognition-Accountability Grid A two-axis chart that places an AI initiative by how much judgement the model performs and how much consequence attaches to being wrong. Decision Head A small service that takes the state of a case and a list of typed questions and returns a probability for each answer, on a model you host. Delegation Cliff The point where pushing work further from human review stops paying for itself. Error Independence The property that makes a second opinion worth having: when one checker is wrong, the other is wrong on different items. Sprawl cost under a fixed monitoring budget What a fleet of agents costs as it grows when the monitoring budget stays flat, which here rises faster than the fleet does. Trial Balance An audit that compares the ledger of what an agent owes with the log of what it ran, and reads only where the two fail to balance. A concept appears here once a public piece carries it. Concepts whose pieces are still queued are not listed. ### Delegation Cliff calculator URL: https://dinand.com/experiments/instruments/delegation-cliff/ Section: Experiments · Instrument Instruments · Delegation Cliff Delegation Cliff calculator Move the twelve organisation levers and read how far the published coefficients say the quality margin moves, and which levers carry the movement. The instrument This instrument needs scripts to draw. Every number it uses is in the table under Source. Each lever starts at zero, meaning no change from your current design. Move a lever by up to one increment either way. Add one level of depth Add depth, buy it back with reviewers Add depth, buy a better model Reset Modelled change in the quality margin 0.00 Every lever is at zero. Each lever's contribution: its change times its coefficient. The pale bar behind each row is what one full increment of that lever is worth. Bars to the right buy margin and bars to the left spend it; the scale is fixed, so a bar's length compares across settings. What it computes The change in the delegation margin, the gap between the quality an organisation of agents delivers and the floor it must hold, when you move one or more design levers away from wherever you start. It shows the net change, each lever's share of it, and how much reviewer capacity would cancel a net loss. The formula in plain words Multiply each lever's change, in increments, by its published coefficient, then add the twelve products. A negative total spends margin and a positive total buys it. This is the additive linear reading of the published coefficients: the article reports descriptive linear slopes and no intercept, so the instrument can give a change from your starting point and never an absolute margin. What this does not say It does not say whether your organisation sits above or below its quality floor, since the article publishes no intercept. It gives no interval, because the article prints none. The slopes describe configurations inside the transition band of the Sobol sample, so they are descriptive and say nothing causal about the space outside the band, and the interactions the article mentions between observability and capacity are absent from an additive reading. Model capability and agent self-check differ by 0.034 and the article reads them as tied. Source Every number comes from Pricing Agent Autonomy (Paper, Medium, 2026-09-03), with some from Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers (Paper, 2026-09-10). The campaign record is Delegation Cliff . Every number this instrument uses Number As published Where Linear-probe coefficients on the delegation margin, Sobol sample, band-only labels: Delegation depth / Reviewer capacity / Task coupling / Model capability / Self-check calibration / Rework cost / Verification depth / Verification coverage / Detection lag / Workload volatility / Trust response / Escalation latency; values: −1.499, 1.456, −0.958, 0.949, 0.915, −0.803, 0.413, 0.308, −0.259, −0.147, 0.102, −0.007; reads: the price of autonomy / the lever that works / second-largest cost / tied for second / tied for second, paid once / friction compounds / complements capacity / complements capacity / slower is worse / minor / no spiral / no detectable effect Paper , The numbers table, and the chart What moves the delegation margin Configurations in the band-only fit 4948 Paper , Chart caption and footer Gap between model capability and self-check calibration, read as tied 0.034 Paper , How to read these, and how not to Unit of one increment, as the lab names it band-standardised increment Paper , Section 1, citing Pricing Agent Autonomy ### SIGIL tier explorer URL: https://dinand.com/experiments/instruments/sigil/ Section: Experiments · Instrument Instruments · SIGIL: does a liability ledger change what an overseer authorizes? SIGIL tier explorer Pick an irreversibility tier and compare what blocking review and review with a liability ledger did to the value of oversight there, beside what the registered prediction said would happen. The instrument This instrument needs scripts to draw. Every number it uses is in the table under Source. Irreversibility tier Show Net value Wrongful blocks The registered prediction What it computes For each of five irreversibility tiers, from invoices that can be undone next cycle to public corrections that cannot be unsaid, the net value per decision of two oversight designs against no oversight at all, the share of decisions each design blocked, the share it blocked wrongly, and the share of worker errors a blocker could have caught. The formula in plain words Net value per decision is the negative of loss plus oversight cost, averaged over the tier, minus the same average with no oversight. Under the convex cost model an uncaught wrongful approval at tier t costs t squared, a wrongful rejection or a wrongful block costs 35% of that, and every consultation of the supervisor costs 0.15. Values are pooled over three seeds. Below zero, oversight costs more than it saves. What this does not say It does not say that review never pays. The result is local to one base model, qwen3.5, sitting in both seats, on synthetic checklist-compliance tasks, with three seeds, and the liability ledger is a thin, simulated form of accountability. Tier 3 produced no worker errors at all, so its value is the consultation fee and nothing more. The post-hoc audit design, which trimmed the worst losses, is outside this view. Source Every number comes from Accountability Makes Oversight Worse: A Pre-Registered Test of Liability-Exposed AI Supervision Across Irreversibility Tiers (Paper, Medium, 2026-09-10), with some from Skin in the Game Made the AI Supervisor Worse (Executive piece, 2026-09-09). The campaign record is SIGIL: does a liability ledger change what an overseer authorizes? . Every number this instrument uses Number As published Where Task family at each tier Invoice approval, reversible next cycle / Refund release, recoverable with effort / Contract clause acceptance, binds until renegotiated / Data deletion, recoverable only from backup if at all / Public correction, cannot be unsaid Paper , Section 3.1, The delegation game Net value per decision against no oversight, tiers 1 to 5 review: −0.13, −0.17, −0.18, −0.17, −2.64; ledger: −0.14, −0.20, −0.16, −0.51, −3.81 Executive piece , The one table to keep (convex cost model, pooled over seeds) Share of decisions blocked and wrongly blocked, tiers 1 to 5, in percent review block: 7.2, 21.6, 2.5, 3.5, 30.5; review wrong: 3.3, 11.9, 1.1, 1.5, 30.1; ledger block: 3.9, 4.3, 1.6, 6.0, 43.4; ledger wrong: 1.5, 2.5, 0.3, 5.9, 42.7 Paper , Table 2, supervisor blocking per tier, convex campaign Worker errors a blocker could catch (wrongful approvals), tiers 1 to 5; tier 3 had no errors share: 93, 56, , 17, 1.4; approves: 157, 115, 0, 17, 1; rejects: 12, 89, 0, 81, 68 Paper , Figure 2, false-approve share of autonomous worker errors, convex campaign Registered series that showed the predicted zero crossing 0 of 12 Paper , Abstract and section 4.1 Second difference at tier 4, review and ledger (the prediction needed it positive) review: −2.49; ledger: −2.95 Paper , Section 4.1 Tier-5 value under the linear cost model, review and ledger, pooled review: −0.656; ledger: −0.889 Paper , Section 4.1, linear campaign Tier 5: review's mean net value against an always-allow replay of the same episodes actual: −3.47; allow: −0.91 Paper , Section 4.2 Registered band for the unsupervised worker's error rate, and the rate the campaigns ran at (convex, linear), in percent low: 10; high: 35; convex: 14.4; linear: 14.5 Paper , Section 2 and section 3.5, the smoke gate What the registered prediction said at each tier, and how it fared tier1: Predicted: review loses money on reversible decisions. Held: review sits at −0.13 and the ledger at −0.14, and every seed shows the loss. It is the only clause of the prediction the data grants.; tier2: Predicted: value climbs toward a zero crossing somewhere along the gradient. Refuted: both designs stay below zero, and the ledger loses more than plain review.; tier3: Predicted: value climbs toward a zero crossing. Refuted, and this tier cannot speak to it: the worker made zero errors here across 6, 000 episodes, so what remains is the consultation fee.; tier4: Predicted: past the crossing, value rises convexly. Refuted: no series crossed zero, and the second difference at tier 4 is −2.49 for review and −2.95 for the ledger, curving down.; tier5: Predicted: the largest gain of all, at the most irreversible decisions. Refuted: value collapses to −2.64 for review and −3.81 for the ledger, the same sign in every seed and under both cost models. Paper , Section 1 and section 4.1, hypothesis H5 ### Trial Balance budget curve URL: https://dinand.com/experiments/instruments/trial-balance/ Section: Experiments · Instrument Instruments · A Trial Balance for Agent Omissions Trial Balance budget curve Set a reading budget and a diary length and read how many skipped steps each of four audits finds, then switch fault classes on and off to see what a full read of the diary finds and misses. The instrument This instrument needs scripts to draw. Every number it uses is in the table under Source. Reading budget 12,800 Diary length, requisitions Audits What a full read finds, by fault class Fault classes Share of injected faults named correctly at a full read (budget 12,800), all diary lengths pooled, 84 faults per class. The balance check is blind to a wrong amount carried on both sides of the book. What it computes The share of injected omissions each audit named correctly, at each of seven reading budgets and seven diary lengths, for one auditor model. A second chart gives detection at a full read by fault class for three of the audits. The formula in plain words Strict detection is the share of injected faults where the audit flagged a fault, named the right class, the right requisition and the right step. Omissions pool the silent skip and the skip the diary claims was done. Each point is 24 audits on the two primary synthetic seeds; each full-read bar is 84. The budget is the nominal reading allowance; the measured prompt tokens at the two ends of the curve are printed where the pieces give them. What this does not say It does not price the commitment ledgers themselves: imbalance-first's score is mostly the auditor copying the imbalanced line into its verdict, and building the two ledgers is outside the measurement. Every number is one auditor, qwen3.5, under schema-constrained decoding with thinking off, on one procurement workflow with one fault per diary. The balance check finds none of the wrong amounts carried on both sides of the book, by construction. Source Every number comes from A Trial Balance for Agent Omissions (Paper, Medium, 2026-10-01), with some from An AI agent can skip a step without leaving a trace of the omission. (Executive piece, 2026-10-01). The campaign record is A Trial Balance for Agent Omissions . Every number this instrument uses Number As published Where Nominal reading budgets 200, 400, 800, 1600, 3200, 6400, 12800 Paper , Section 3, Design, and the columns of Figure 3 (six doublings from 200 to 12,800) Diary lengths, in requisitions 1, 2, 3, 5, 8, 12, 16 Paper , Section 3, Design Omission detection per arm; one row of seven budgets per diary length, rows in the order of the lengths imbalance: 0.96, 0.96, 1.00, 1.00, 1.00, 1.00, 1.00 / 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00 / 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00 / 1.00, 0.96, 1.00, 1.00, 1.00, 1.00, 1.00 / 1.00, 1.00, 1.00, 1.00, 0.96, 1.00, 1.00 / 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00 / 0.96, 0.92, 1.00, 1.00, 0.96, 1.00, 1.00; logreview: 0.04, 0.17, 0.29, 0.29, 0.33, 0.25, 0.29 / 0.00, 0.21, 0.21, 0.54, 0.46, 0.50, 0.58 / 0.00, 0.08, 0.08, 0.21, 0.29, 0.38, 0.33 / 0.00, 0.00, 0.00, 0.21, 0.21, 0.42, 0.46 / 0.00, 0.00, 0.08, 0.08, 0.08, 0.17, 0.21 / 0.00, 0.00, 0.04, 0.08, 0.21, 0.25, 0.29 / 0.00, 0.00, 0.00, 0.08, 0.04, 0.00, 0.12; sample: 0.17, 0.54, 0.25, 0.25, 0.25, 0.25, 0.25 / 0.04, 0.12, 0.46, 0.50, 0.39, 0.42, 0.50 / 0.00, 0.04, 0.12, 0.21, 0.33, 0.29, 0.29 / 0.04, 0.04, 0.00, 0.17, 0.29, 0.33, 0.29 / 0.00, 0.00, 0.00, 0.00, 0.00, 0.17, 0.21 / 0.00, 0.00, 0.00, 0.08, 0.04, 0.08, 0.25 / 0.00, 0.00, 0.00, 0.00, 0.00, 0.08, 0.21; checklist: , 0.25, , 0.50, , , 0.58 / , 0.08, , 0.62, , , 0.58 / , 0.00, , 0.33, , , 0.50 / , 0.00, , 0.33, , , 0.71 / , 0.04, , 0.25, , , 0.58 / , 0.00, , 0.08, , , 0.46 / , 0.00, , 0.04, , , 0.38 Paper , Figure 3, strict omission detection by trace length and budget, synthetic seeds 1 and 2, n = 24 per cell Detection threshold that defines C80 0.80 Paper , Section 4, H1 Measured prompt tokens for imbalance-first at the smallest budget, per diary length 563, 586, 611, 666, 745, 850, 954 Paper , Figure 2, imbalance-first C80 in measured prompt tokens per diary length (at budget 200) Measured prompt tokens for log review at a full read, per diary length 994, 1683, 2421, 3882, 6173, 9044, 12184 Executive piece , Figure 2, reading cost per audited episode Correct omission findings per thousand prompt tokens at 8 and 16 requisitions: imbalance-first, log review at its best budget, log review at a full read k8: 1.34, 0.063, 0.034; k16: 1.00, 0.035, 0.010 Paper , Section 4, H1, the finite descriptive Log review reading the whole diary at about 6,200 tokens, pooled over 1 and 8 requisitions 0.25 Executive piece , Headline figures Detection at a full read by fault class classes: Silent omission / Phantom completion / Value commission / Consistent commission; imbalance: 1.00, 1.00, 1.00, 0.00; checklist: 0.726, 0.357, 0.976, 0.952; logreview: 0.476, 0.179, 0.976, 1.000 Paper , Figure 4, strict detection at full read by fault class, synthetic seeds 1 and 2, n = 84 per class ### The protocol URL: https://dinand.com/research/methods/ Section: Research Methods The protocol The lab studies how organisations delegate work to AI agents and how they oversee it. Every campaign follows the same eight steps, and the steps are built so the lab cannot quietly steer a result toward the answer it hoped for. Ask one question A campaign starts with a question a buyer or a regulator would recognise, such as whether a second model checks the first or what a liability ledger does to a reviewer. It names the quantity that would answer it and the null result that would mean the idea adds nothing. Register the prediction, the bands and the kill conditions Before any unit runs, the lab writes down the hypotheses, the band each number has to land in to count as a hit, the statistics that will be used and the conditions that stop the run. The document is dated and frozen. A later change goes in as a dated amendment that says what moved and why, and it is written before the data it affects is read. Each campaign tests its instrument first. A small calibration stage checks that the task is neither trivial nor impossible and that the harness records what it should. That gate is allowed to test whether the instrument works. It is never allowed to require the result the campaign exists to measure; the lab learned that by halting a campaign on exactly that mistake. Attack the design A reviewer whose job is to break the headline goes over the design and the analysis: a confound, a statistic that hides a split, a ceiling or a floor, a scoring rule that favours one arm. What the review finds is fixed before the run or stated as a limit in the piece. Run to the registered rule The campaign runs on open models on hardware the lab owns, so any run can be repeated. It stops where the registration says it stops. A failed gate, an unmet validity check or a spent compute budget ends the run, and what is complete at that point is analysed as it stands. The lab does not extend a run, loosen a threshold or add a round to rescue it. Verify from the raw files A separate pass, with fresh code and no access to the original analysis, recomputes the reported numbers from the raw records. Each piece says how many numbers were checked and how many matched. Where the two passes disagree, the disagreement is chased down and the piece carries the verified value. On the sprawl-cost campaign this pass found duplicate runs and moved a verdict from supported to refuted before publication. Publish the negative results A refuted hypothesis is published with the same care as a confirmed one, next to the registered prediction it contradicted. A null result is reported as a bound, with the size of effect the study could have detected. What was measured is kept apart from what the lab infers, and an explanation is labelled as one until a test has been run against it. Release the code Each campaign keeps its registration, its harness, its unmodified ledger and its analysis together, so the published numbers can be traced to the rows that produced them. The code and records are released with the pieces as each campaign is cleared for it. Log every correction When a published number turns out to be wrong, the correction is made in the open and dated, and the original number stays visible beside it. The measured claims index keeps a row for every headline number, so a correction has a fixed place to land. ## Interactive pieces Self-contained pages (films, experiments and side projects). Some need JavaScript or WebGL to run; their written content is reproduced below. ### The Sympathy of Clocks URL: https://dinand.com/reading-notes/the-sympathy-of-clocks/ Section: Reading Notes · Film · Film · 6:20 · Essay and experiment What a feverish Dutch scientist saw in his sickroom in 1665, and why it took three centuries for anyone to build on it. I made it to open an AI keynote. Watch it full screen with the sound up. Page text: The Sympathy of Clocks | Dinand Tinholt The Sympathy of Clocks Film Essay Experiment Timeline ← dinand.com A short film and an essay The Sympathy of Clocks What a feverish Dutch scientist saw in his sickroom in 1665, and why it took three centuries for anyone to build on it. 1665 the observation 1975 the mathematics 6:20 running time loading the film Play the film · sound on ▶ 0:00 / 6:20 CC ⤢ Made to open an AI keynote. Watch it full screen with the sound up. Space to pause · F for full screen · C for subtitles The essay A fever in The Hague In February 1665, Christiaan Huygens lay sick in bed in The Hague. At thirty-five he had already worked out that Saturn wears a ring, and his pendulum clock kept better time than any machine built before it. Two of those clocks hung from a single wooden beam in his room, prototypes for a trial at sea. A fever leaves a man little to do but look at things, so for days he watched them. The pendulums had started out of step and ended in perfect opposition, each swinging left as the other swung right. When Huygens disturbed one, the pair found each other again inside half an hour. Moving the clocks to separate supports broke the spell. In a letter to his father he called it a sympathy of clocks and blamed vibrations in the beam too faint for the eye to catch. He reported it to the Royal Society that March and went back to optics. 1 Caspar Netscher · 1671 Christiaan Huygens Painted six years after the winter he spent watching two of his own clocks agree on a rhythm nobody had set. Historians count this as the first recorded case of spontaneous synchronization, two coupled systems falling into order with nobody in charge. Complexity science still uses it as a founding example. Huygens knew more about precision machinery than anyone alive, and he wrote the whole thing off as a curiosity. The mathematics that explains his clocks took until 1975 to arrive. The shelf was already stocked By the 1660s most of what an AI course opens with sat in plain view. Hobbes wrote in Leviathan , in 1651, that reasoning is “nothing but reckoning,” the adding and subtracting of thoughts, and every educated reader in Europe had an opinion about the book. Pascal was nineteen when he built a calculating machine in 1642. A gambling dispute led Pascal and Fermat to probability theory in 1654, and Huygens published its first textbook three years later. In 1662 John Graunt, a London haberdasher, went through the city's weekly death registers looking for patterns and founded statistics in the process. 1651 Thomas Hobbes Argued that thinking is a kind of arithmetic. 1642 Blaise Pascal Built a brass box that could add. 1657 Christiaan Huygens Wrote the first textbook on chance, two years before his book on Saturn. 1662 John Graunt Counted London's dead and found patterns in the totals. Leibniz came closest to putting the pieces together. In 1673 he showed the Royal Society his stepped reckoner, the first machine that could multiply. He spent much of his life on a universal symbolic language in which two philosophers could settle any dispute by sitting down and saying calculemus , let us calculate. In November 1676 he travelled to The Hague to spend several days talking with Spinoza, who had already folded the mind into nature and treated mental life as one more aspect of the substance that makes up matter. 2 You can follow the paper trail from that visit all the way to a transformer model. Gottfried Wilhelm Leibniz · 1673 The Stepped Reckoner It multiplied with a hand crank and jammed often enough that Leibniz kept tinkering with it for decades. Why the clock won The century's own success slowed everything down. Calculus arrived in the 1670s and 1680s and cracked planetary orbits and the flight of a cannonball. Problems like these are linear or close to it. The whole behaves like the sum of its parts, and a patient person with pencil and paper can solve them. After a generation of such wins, Europe came to picture the universe as a clock, wound once and running on rules anyone could write down. Much of nature behaves otherwise. Its parts feed back on one another and change each other's rules as they go. Huygens' two clocks were a system of that kind, which is why his own equations could not touch them. Nonlinear problems give way to iteration: compute a step, feed the answer back in, and repeat it a million times. Nobody in 1665 owned a machine with that kind of patience, so the problem sat. Poincaré breaks the clockwork In 1889 Henri Poincaré won a prize set by the King of Sweden for work on whether the solar system is stable. Even three bodies under gravity, he showed, admit no general formula, and in the tangle of their orbits he found behavior so intricate he declined to try drawing it. In 1961 the meteorologist Edward Lorenz restarted a weather simulation from a printout and typed 0.506 where the machine had stored 0.506127. Within a couple of simulated months the new weather had nothing in common with the old. His 1963 paper on sensitive dependence on initial conditions founded modern chaos theory. 3 A burst of results followed in the 1970s. Robert May showed in 1976 that the simplest population equation a biology student meets tips from steady numbers into full chaos as one parameter is nudged upward. Mitchell Feigenbaum found that the tipping follows the same ratio, 4.669 and change, in systems with nothing else in common. Mandelbrot named the geometry in 1975: fractals. That same year Yoshiki Kuramoto published the equation Huygens never had, a solvable model of coupled oscillators that shows the exact moment a crowd of independent tickers snaps into shared time. The experiment Turn up the coupling Every light in the scene keeps its own natural rhythm, the way a firefly or a pendulum clock does, and drifts a little toward the rhythm of the crowd. The slider sets how hard that pull is. At low coupling the lights flicker at random. Push the slider past 0.8 and the crowd tips into unison while the coherence r climbs toward one. Kuramoto solved this transition exactly in 1975. dθ i /dt = ω i + K · r · sin(ψ − θ i ) Coupling K 0.30 Scatter the phases Jump past the threshold coherence r 0.00 What the clocks were trying to say With Kuramoto's result in hand, the same machinery turns up everywhere. Along riverbanks in Thailand, thousands of fireflies flash in unison, each one obeying a rule short enough to fit in a sentence: if a neighbor flashes first, flash a little sooner next time. Pacemaker cells in the heart keep time the same way, and so do the neurons whose rhythms let you read this page. Scientists call this emergence. Simple rules and dense coupling at the bottom produce capable behavior at the top, and nothing anywhere holds a blueprint of the whole. In 1943 McCulloch and Pitts showed that networks of idealized neurons can compute anything logic can express. Rosenblatt's perceptron was learning from examples by 1958, and backpropagation made deep networks trainable in the 1980s. By the 1990s the theory was in decent shape and still waiting, as Huygens had, for a machine patient enough to run it at scale. A fever leaves a man little to do but look at things. The Hague, 1665 The patience machine finally shows up ENIAC, unveiled in 1946, did 357 multiplications a second with eighteen thousand vacuum tubes and drew as much power as a small neighborhood. The transistor came a year later. Integrated circuits followed in the late 1950s, a whole processor fit on one chip by 1971, and for fifty years the number of transistors on a chip doubled about every two years. Ballistic Research Laboratory ENIAC Glen Beck and Betty Snyder at the machine that made a few hundred multiplications a second feel like the future. Video games supplied the next turn. Drawing a 3D scene sixty times a second means recalculating millions of points and pixels in parallel, so graphics cards grew into machines built to multiply large matrices fast. Nvidia opened that hardware to general computing in 2007. Five years later two Toronto students trained a deep neural network on a pair of consumer gaming cards and won the ImageNet vision competition by a margin nobody in the field had seen before. 4 The transformer arrived in 2017. Under the vocabulary it is matrix multiplication arranged so that every token in a sequence can attend to every other, and graphics silicon runs that arrangement well. Architecture and hardware locked together the way Huygens' pendulums did. One modern AI chip now performs about two thousand trillion low-precision multiplications every second. Out of multiplication at that scale came the emergence this whole lineage had been pointing toward, systems that write and plan even though no single part of them understands anything. Filing things under curiosity has a price I open AI sessions with this story because of how long it took to end. The best-equipped mind in Europe watched something new happen on his own wall, gave it a lovely name and went back to problems he could solve. The observation was sound, and Huygens filed it under curiosity, where it stayed for three centuries. Plenty of companies are running the 1665 experiment again right now. A working demonstration of machine intelligence hangs on the beam, people gather round and agree it is fascinating, and the pilot gets a lovely name. The three hundred years live in the distance between that moment and the day the work changes. The tools have caught up with the observation. Huygens saw it from his bed in 1665, centuries before anyone had the equipment to act on it. The timeline From the beam to the chip 1642 Pascal builds a calculating machine He was nineteen. 1651 Hobbes: reasoning is reckoning Leviathan argues that thinking is a kind of arithmetic. 1654 Pascal and Fermat invent probability A gambling dispute turns chance into mathematics. 1662 Graunt founds statistics London's death registers become data. 1665 Huygens sees the sympathy of clocks Two pendulums on one beam fall into step, and he files it as a curiosity. 1673 Leibniz demonstrates the stepped reckoner The first machine that multiplies, shown in London. 1676 Leibniz visits Spinoza in The Hague Several days of talk with the philosopher who put the mind inside nature. the trail goes quiet for 214 years 1890 Poincaré breaks the clockwork Three bodies under gravity turn out to have no general solution. 1943 McCulloch and Pitts model the neuron Networks of simple units can compute anything logic can. 1946 ENIAC is unveiled 357 multiplications a second. 1963 Lorenz publishes the butterfly A rounded number sends a simulated sky somewhere new. 1975 Kuramoto solves the clocks An exact model of coupled oscillators, the same year Mandelbrot coins the word fractal. 2012 Deep learning on two gaming cards AlexNet wins ImageNet and the scaling era begins. 2017 The transformer Attention built from matrix multiplication, which graphics chips run fast. Now Emergence on purpose Models whose abilities come from billions of simple interactions. Notes Huygens described the effect in letters to his father in February 1665 and in a report read to the Royal Society that March. Bennett, Schatz, Rockwood and Wiesenfeld published a modern analysis of the two-clock system in 2002 and confirmed his explanation about vibrations in the beam. ↑ Spinoza's Ethics circulated in manuscript during his lifetime and appeared in print in 1677, months after his death. Antonio Damasio's Looking for Spinoza (2003) makes the case for him as an ancestor of modern neuroscience. ↑ Edward Lorenz, “Deterministic Nonperiodic Flow,” Journal of the Atmospheric Sciences , 1963. The butterfly entered through the title of his 1972 talk: “Does the Flap of a Butterfly's Wings in Brazil Set Off a Tornado in Texas?” ↑ Krizhevsky, Sutskever and Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” 2012. AlexNet trained on two Nvidia GTX 580 gaming cards and reached a top-5 error of about 15 percent, against 26 percent for the runner-up. ↑ Film narration Kokoro neural voice (af_heart) Score Original, composed for the film Christiaan Huygens Caspar Netscher, 1671 · public domain The Philosopher in Meditation Engraving after Rembrandt · public domain View of The Hague Rijksmuseum SK-A-4870 · CC0 London's Dreadful Visitation Bills of Mortality, 1665 · public domain Huygens clock replication H. M. Oliveira & L. V. Melo · CC BY 4.0 Leviathan frontispiece Abraham Bosse, 1651 · British Library · CC0 The Pascaline Photo by Rama · CC BY-SA 3.0 FR Leibniz, Spinoza, Poincaré Portraits · public domain Stepped Reckoner, ENIAC Public domain Systema Saturnium, pendulum clock Christiaan Huygens · public domain Dinand Tinholt The Sympathy of Clocks · 2026 ### Many Worlds URL: https://dinand.com/experiments/many-worlds/ Section: Experiments · Simulation · Scenario planning Plays a business decision out across hundreds of thousands of simulated futures, then tests every possible response against the same draws and recommends the one that holds up. Five library scenarios with illustrative figures. Change any assumption in the side panel and it reruns in your browser. Page text: Many Worlds | Dinand Tinholt Many Worlds Scenario planning across a million futures ← dinand.com Hide introduction About this tool Scenario planning that tests every response against the same futures Many Worlds plays a business scenario out hundreds of thousands of times. Each simulated future draws its own mix of demand, prices, costs and shocks, so the spread of outcomes shows how wide the plausible range runs before anyone commits money. The engine then runs every possible response through all of those futures. It compares each set of moves against doing nothing on identical draws and recommends the one that holds up best for the risk appetite set in the side panel. The five library scenarios carry illustrative figures. For a client, replace them with real numbers under Assumptions or paste historical data to fit the ranges. Recommended response Doing nothing Each colour keeps its meaning in every chart on the page. Pick a scenario Choose one from the library on the left, or describe a client situation in plain words. The engine plays out the futures Every driver is drawn from its range, and shocks land at random times with random force. Every response meets the same futures Each combination of moves is scored on identical draws, including moves held in reserve until trouble starts. Any gap between two policies therefore comes from the moves themselves. The call gets stress-tested The winner is rerun with one assumption changed at a time before it reaches the top of the page. Policy A combination of moves, where any move can start now or wait in reserve as an option. Option A move prepared now for 15% of its cost and exercised only if its trigger fires. Worst tenth The average result across the worst 10% of futures, a gauge of how bad a bad outcome gets. Risk-adjusted score Expected value blended with the worst tenth, weighted by the risk posture you choose. Unacceptable outcome A result below the floor set in the side panel, which by default is the level doing nothing misses in a fifth of futures. Starting 0 policy-futures simulated Starting the engine. Scenario Describe the company and what worries it. Mention any responses already on the table. Draft with Claude Stop Save Remove from library Saved scenarios are visible to everyone this page is shared with. Decision settings Risk posture Growth Balanced Defensive Horizon 3 years 5 years Upfront budget Policies that cost more upfront are ruled out. Unacceptable below Futures per finalist Draw a fresh set of futures About Decision Plan Moves Outcomes Stress tests Policies Ask Claude Assumptions Recommended response Refining Running the first simulation. Results CSV Decision memo Reasoning Why this is the call Robustness How sturdy the call is Plan The plan, quarter by quarter Committed moves start this quarter, while an option gets prepared now and is exercised only if its trigger fires. Moves Value of each move Each bar shows how far the risk-adjusted score moves when you add a move to the recommended set or take it out. Every comparison runs on the same futures. Outcomes Where the outcomes land Failure modes When the call is wrong The engine searches the futures for the conditions that best explain where the recommendation loses. Stress tests How the call fares when assumptions are wrong Exposure Exposure that remains The average outcome under the recommended response when each uncertainty sits at either end of its range. Budget Return on more budget The best score the search found at each level of upfront spend. Search Every policy tested Each dot is one combination of moves, committed or held as options. A dot sits higher when its expected value is larger and further right when its worst tenth of futures hurts less. Finalists The finalists, side by side Analyst Ask Claude Claude reads the scenario and the results, and reruns the simulation itself to test a what-if. It uses your own Claude account. Write the board brief Short sections on the call and its support, what would change it and the first 90 days. Argue against it Claude makes the strongest case against the call and names the assumptions that would flip it. Ask Stop Inputs Assumptions The engine uses every input listed here, and editing any value reruns the simulation. Each range runs from a plausible low to a plausible high, with its peak at the most likely value. Fit ranges and correlations from data Paste a table with a header row and one column per driver, in the driver's own unit such as yearly % change. Each fitted range spans the 5th to the 95th percentile and peaks at the median. Correlations come from the matched columns. Read columns Method. Each future draws a value for every driver from its triangular range, linked through a Gaussian copula where correlations are set. It also draws when each event strikes and how hard, along with quarterly demand noise and how well each move performs. Because every policy is scored on the same draws, the gaps between policies are measured far more precisely than the outcomes themselves. The engine first ranks every combination of committed moves. A local search then tests whether each move does better held as an option, and the seven best policies go on to the full run. Stress tests rerun those finalists with one assumption changed at a time. Limits. The library figures are illustrative. An option fires on a fixed trigger, which is the first quarter its event starts or its one-off shock lands on the bad side. Preparing an option costs 15% of the move, and exercising it carries a 15% rush premium. Effects combine multiplicatively, and competitors in the model never respond to the company's moves. ### Dead Reckoning URL: https://dinand.com/experiments/dead-reckoning/ Section: Experiments · Field guide · Physical AI and world models A world model is dead reckoning for machines. A small ship plans by imagining forty futures at every fix, then drifts on a current its model has never heard of. Scroll to follow the vessel, then take the helm yourself. Best on a laptop. Page text: Dead Reckoning | Dinand Tinholt DR 0% Dead Reckoning ← dinand.com The loop Arithmetic The stack Deployment The field Readiness Pilot plan Terms Reading Take the helm Arrivals 0 Groundings 0 Position error 0.00 nm Since last fix 0 min A field guide to physical AI and world models A world model is dead reckoning for machines. Before satellites, a navigator worked out a ship's position from heading, speed and the hours since the last landmark. Robots still do it. A world model extends the habit from position to everything else in the scene. Scroll to follow the vessel On the chart · Fan of tracks Rollouts At every fix the vessel imagines forty futures. Each pink line is one of them, played forward inside the model without the vessel moving at all. Planning is choosing among them, and the solid line is the choice. This view: fix every 3 min, imagines 6 nm ahead On the chart · Current Unmodelled dynamics The arrows are a current the model has never heard of. In a robot it is friction, cable slack, a box heavier than yesterday's. The plan assumes still water, so the vessel drifts off the line it believes it is sailing. This view: current raised to 3 knots On the chart · Circle of uncertainty Compounding error The ring marks where the model thinks the vessel is, and how wrong that could be. It widens with every minute since the last fix, because each predicted step inherits the error of the one before. This view: fix every 14 min On the chart · Fix Observation A fix is a look at the world: a camera frame, a force reading. The ring collapses, the model replans from the truth, and the drift starts again. The more often a robot can afford to look, the worse its model is allowed to be. This view: fix every 4 min in a 2.5 knot current On the chart · Grounding Failure in the world This crew checks its position every thirty minutes in a four-knot current. The plan stays sound inside the model while the vessel runs onto the rocks. Stay here a moment and watch the groundings counter. This view: fix every 30 min, 4 knot current Your turn Take the helm. Click the water to set a destination. Drag to turn the chart. Attentive Rarely looks up Perfect model Fixes come often and the model sees far ahead. The current still pushes, and each fix absorbs it. Takes a fix every 3 min Imagines ahead 5 nm Unknown current 1.5 kn Track sailed Futures imagined Future chosen Believed position The arithmetic of a long task A demo is a short rollout. A shift is a long one. Set how often one step works and how many steps run without a check. Each row below is one simulated run, and it stops at its first failure. 37% of runs finish clean Each step works 99% Steps in a row without a check 100 Each mark is one step. Pink marks the first failure. Why now Language models learned from text people had already written. Physical AI has no such archive. The record of what happens when a gripper closes on a wet glass has to be collected by robots, or imagined by a model good enough to stand in for them. World models moved to the centre of the field because they give a machine a cheap place to practise. The stack A physical AI system has to see, predict, act, practise and prove it. Five layers, from the sensors up. Select a layer to read what it does and what to ask about it. Layer 2 of 5 Predict ↓ ↑ For your next vendor meeting Five questions to put to anyone selling physical AI. Copy the questions Which senses does the policy use at run time, and how often does it read them? Does the system predict consequences before it moves, and how many steps ahead? What share of the training data came from robots like the one on offer? How much of the training was real, and how was the transfer from simulation measured? What was the success rate over the last thousand unattended cycles, and who counted? Deployment The charted water is narrower than the demos suggest. For a decade the pattern has held: deployment arrives first where the environment is bounded and a mistake can be retried. Charted In commercial service Driverless ride-hailing Paying passengers in more than a dozen US cities, inside mapped service areas with remote support on call. Warehouse picking and sorting Arms pulling mixed items from bins across full shifts. The task is bounded and a miss gets retried. Visual inspection Learned models checking welds, labels and surfaces. Nothing moves on the model's say-so, which keeps errors cheap. 4 7 5 Coastal Paid pilots on narrow tasks Humanoids on factory floors Moving totes and loading fixtures in a few plants and warehouses, supervised, and slower than the people beside them. Driverless freight Trucks on fixed highway routes between depots, in fair weather. Mobile arms in hospitals and labs Fetching and delivering along corridors the robot already knows. 38 71 54 63 Open water Still research General home robots Every kitchen is a new environment. Many of the impressive demonstrations have a person at the controls. Soft and wet things Cloth, cables, food. Touch sensing and simulation both fall short here. Hours without supervision Error compounds over a long task, as the arithmetic above shows. The placement is a judgment call, current to October 2026. The field Thirty-six years from an idea to an $8.2 billion acquisition. Each mark is a system or a moment worth knowing. Select one to read about it. The 2025 column is where the large labs committed. Readiness Read the barometer before you commit to a pilot. Pick one physical task and answer for that task alone. The answers loaded here are an example for a parcel-sorting line. Replace them with yours. What holds you back first Copy this reading Showing example answers. Passage plan Sail a first pilot in legs, and leave each one on evidence. From harbour to open water. Write each exit condition down before the leg starts. Leg 1 Record Instrument the station. Capture video and machine actions together, with a pass or fail on every cycle. Leave when you hold a few hundred labelled cycles that cover the ordinary variation. Leg 2 Rehearse Adapt a policy to the task and run it on a test rig or a simulated copy of the station, away from production. Leave when more data stops improving the success rate on the rig. Leg 3 Shadow The model watches live production and proposes actions that nobody executes. Compare each proposal with what the operator did. Leave when its misses are few and every one of them has an explanation. Leg 4 Supervise The machine acts while a person stands by with a stop button. Count every intervention. Leave when interventions per shift fall below the number you wrote down before starting. Leg 5 Release The machine works alone inside fixed limits, with an automatic check at the point of work. Then widen one variable at a time and repeat the count. Terms If a term cannot be unpacked, it is decoration. Nine words that come up in every conversation about physical AI, one layer down. World model A learned function that takes the current state of an environment and an action, and returns the next state. A system that cannot be asked "what happens if I do this?" is a video generator. Vision-language-action model A policy that reads camera frames and a written instruction and writes motor commands. The language part usually comes from a pretrained model, and the action part is trained on robot demonstrations. Rollout A sequence of predicted steps, each fed from the one before. A planner scores many rollouts and executes the opening moves of the best one. The fan on the chart is forty of them. Model predictive control Plan over a horizon, act on the first part of the plan, observe, plan again. The vessel on the chart does exactly this, and so do most robots that use a world model at run time. Latent space The compressed internal description a model keeps of a scene. Predicting there is cheaper than predicting pixels, at the cost of being harder for a person to inspect. Sim-to-real gap The drop in performance when a policy trained in simulation meets real surfaces, lighting and wear. Teams narrow it by randomising the simulation and by mixing in real data. Teleoperation A person drives the robot through a task while every frame and joint command is recorded. Most action-model training data is still made this way, one demonstration at a time. Cross-embodiment Training one policy on data from many robot bodies so that skills learned on an arm with two fingers carry over to a humanoid hand. Digital twin A simulated copy of a specific site, kept in step with the real one. A generic simulation teaches a skill, and a twin lets you rehearse it on your own floor plan. Reading Sources worth an evening. Primary sources, from the paper that named the idea to last week's acquisition. World Models Ha and Schmidhuber, 2018. An agent learns to drive inside its own dream. Interactive, and still the clearest introduction. A Path Towards Autonomous Machine Intelligence LeCun, 2022. The argument for predicting in representation space. Mastering Diverse Domains through World Models Hafner and colleagues, 2023. Dreamer V3. From Words to Worlds Fei-Fei Li, 2025. The case for spatial intelligence as the next frontier. Genie 3: A New Frontier for World Models Google DeepMind, August 2025. Interactive worlds at 24 frames a second, consistent for a few minutes. V-JEPA 2 Meta, 2025. Video pretraining followed by robot planning without task-specific training. Cosmos NVIDIA. Open world foundation models for robotics and driving. π0 Physical Intelligence, 2024. A general robot policy trained across many robot bodies. Atlas World Labs, 2026. Their world model for spatial intelligence. AMD to Acquire World Labs AMD, September 2026. The announcement and its stated reasoning. Dead Reckoning A field guide to physical AI and world models. Corrected to October 2026. Not to be used for navigation. ### Rulebook URL: https://dinand.com/experiments/rulebook/ Section: Experiments · Live model · Reinforcement fine-tuning A customer reply policy written as checks a machine can run, then used as the reward to fine-tune a model. A small transformer learns it live in your browser tab. Training starts by itself and runs for a few minutes. Larkfield, the retailer in the examples, is fictional. Page text: Rulebook | Dinand Tinholt Rulebook ← dinand.com Live training Check a reply Audit past replies Run it on Tinker A working demonstration of reinforcement fine-tuning Teach a model your customer reply policy Any company that answers customers at volume keeps a rulebook for what its agents may write. Refunds have thresholds and some promises are off limits. A general-purpose model sees that rulebook only if the whole policy rides along in every prompt, on every call. This page trains the policy into the model itself. A small model learns it live in your browser tab, through the same loop that Thinking Machines' Tinker service runs on its open-weights Inkling-Small. Larkfield, the retailer in the examples, is fictional. Policy Nine banned habits, three required elements and a refund rule, each written as a check a machine can run. Score The model drafts replies to customer cases. The checks turn each draft into a reward between 0 and 1. Update Drafts that beat their batch average pull a small adapter toward their wording. Repeat 160 rounds of 16 drafts. The model's original weights never change. Fully compliant replies 8% Right refund decision 70% Replies scored 0 step 0 Run training Start over Each square is one reply the model wrote, scored the moment it was written. Each column is one training step, read left to right. Compliant Fails a wording or required check Wrong refund decision sample forward_backward optim_step save_state Check any reply against the policy Paste a reply to see which checks it fails and the reward the training loop would give it. The same scorer runs the live training and the Tinker kit. Audit the replies your team already sent Paste past replies to see which pass. The compliant ones become a training file, so a first supervised round can teach the model your team's own wording. Run it on Inkling-Small with Tinker The kit runs this loop on Thinking Machines' open-weights model from a laptop, while Tinker supplies the GPUs. You keep the trained adapter. 1 python evaluate.py Measure base Inkling-Small on fresh cases, with and without the policy in its prompt, before any training. 2 python train.py Fine-tune a LoRA adapter with the policy as the reward. A default run comes to an estimated $3.56 at today's prices. 3 python evaluate.py --tuned Compare all four arms on trained kinds of case and on one kind training never saw. Within the rule Outside the rule New draft Customer wrote C Policy rule, and where this case falls Base model Frozen weights, never shown the policy Same model, with the adapter Not trained yet Compliance over the run Fully compliant Right decision Stay close to the base model A penalty on drifting from the base model's wording. Without it, the fluency figure above falls step after step. Check the decision against the case facts Switch it off and the adapter learns the wording rules while refund decisions stay where they started. Change a switch, then press Start over. What the adapter changed Copy adapter Words it now avoids Words it now prefers Failure rate by check Base model With the adapter Decision checks What runs here A three-layer transformer with weights, pretrained on synthetic replies for Larkfield, a fictional retailer. It makes the right refund call about seven times in ten because it has never seen the policy. Each step scores 16 sampled replies against the policy and the case facts. The scores update a rank-8 LoRA adapter and a bias on the output layer, as train.py does on Tinker. Plan renewal stays out of training as a held-out test. Five seeded runs of 160 steps 8% → 93% Fully compliant replies on the case types the adapter trained on. Right refund decisions went from 70% to 95%. Each run passed 80% compliant, as a ten-step average, by step 57. Refund calls came later Wording failures fell from the first step. Refund decisions sat near 70% for about 30 steps, then climbed. Renewals got worse Plan renewals stayed out of training. Right decisions on them fell from 75% to 67%, since nothing taught the adapter their 14-day window. Unchecked rules stay unlearned With the decision check off, the training reward reached 0.98. Refund decisions stayed at the base model's 70%. Drift needs a brake Without the base-model penalty, the base model's score for the replies fell from −0.18 to −0.48 per word over 160 steps. Reply Weak example Compliant example Under the policy, the remedy is allowed not allowed not stated Verdict Reward the trainer would receive The policy, check by check Check Kind Penalty What it says Reward is exp(−penalty ÷ 4), scaled down outside 45 to 110 words. A reply with no penalty inside the word band counts as fully compliant. Past replies three examples loaded Clear Put a line of three hyphens between replies. Nothing pasted here leaves this browser. Audit Reply Words Verdict Failed checks Training file compliant replies only Copy The replies that pass, in the chat format Tinker's supervised recipe reads. The kit's transcripts.py builds the same file from a folder. Case facts are unknown here, so the refund decision goes unchecked. Four ways to run the same model Prompt carries the case Prompt carries the case and the policy Base Inkling-Small 1 Needs an API key 2 Needs an API key Trained adapter 3 Needs one training run 4 Needs one training run This demonstration policy is 374 Inkling tokens, about 22 cents per thousand replies at $0.58 per million prompt tokens. A real policy manual runs far longer, and the prompt carries all of it on every call. Estimated bill for a default run $3.56 Sampling 5,120 replies, 0.72M tokens at $1.44 $1.03 Training on them, 1.24M tokens at $1.73 $2.14 Reading the cases, 0.52M tokens at $0.58 $0.30 Fluency check, 0.15M tokens at $0.58 $0.09 An estimate for 40 steps of 128 replies at 140 tokens a reply, thinking effort 0, using Inkling-Small's discounted Tinker prices per million tokens on 6 October 2026. List prices are double. One training step in Tinker's calls sampler = training_client.save_weights_and_get_sampling_client() for case in batch: replies = sampler.sample(case.prompt, num_samples=8, ...) rewards = [policy.reward(text(r), case.eligible) for r in replies] mean = sum(rewards) / len(rewards) datums += [make_datum(case.prompt, r, x - mean) for r, x in zip(replies, rewards)] training_client.forward_backward(datums, "importance_sampling") training_client.optim_step(AdamParams(learning_rate=4e-5)) Abridged from train.py in the kit. What the cases leave out Each case gives the model what an agent sees on screen and the customer's message. The policy thresholds stay inside the reward, so a trained model has to find them by trial. The evaluation runs every arm on fresh cases of the trained kinds and on one kind of case that training never saw. Where this applies Any message a company sends in volume under rules a machine can check fits this method. Payment reminders, dispute responses, claim decisions and supplier letters all qualify. The policy file is the part you replace. Most checks are patterns with a penalty attached. The refund rule compares the decision in a reply with the facts of the case. Anything that resists being written as a check stays with a human reviewer or the optional grading model in the kit. How far the evidence goes The small model's pretraining and seeded runs ran here, along with tests that match the browser engine to a numpy reference and the Python scorer to the JavaScript one. The Tinker kit ran against stand-in clients built on the real Tinker SDK types and the real Inkling tokenizer. No run has touched the live Tinker service yet, so this page holds no Inkling weights or Inkling replies. The small model is about a millionth the size of Inkling-Small and learned from a corpus written so that a compliant wording always exists. Rulebook is an independent demonstration of fine-tuning with Thinking Machines Lab's Tinker API and Inkling-Small. Thinking Machines Lab has not endorsed it and has no affiliation with it. Larkfield and its policy are fictional. ### The Meridian Library URL: https://dinand.com/side-projects/meridian-library/ Section: Side Projects · 3D room · A library to wander A two-storey reading room you can walk through. Pull a book from the shelf and it comes to you; turn its pages, light the fire, let it rain, or climb the spiral stair to the gallery. Best on a laptop with the sound on. Books you keep and notes you write stay in your own browser. ### Breslau URL: https://dinand.com/side-projects/breslau/ Section: Side Projects · Thought experiment · Insurance for AI agents An insurer designed from a blank sheet for AI agents that move money. A signed meter watches every tool call, and risky payments wait for a person. Breslau holds no insurance license. Every rate and figure on the page is illustrative. Page text: Breslau | Dinand Tinholt Breslau is a thought experiment drawn up in October 2026. Nobody can buy cover from it, since it holds no insurance license. DINAND TINHOLT Side Projects / Breslau How it works Rates Risk Economics Software Plan Company Open the live desk A thought experiment in insurance for AI agents Authorization and insurance for every payment your agent makes. Breslau watches every tool call an agent makes, much as a card network watches each card payment. Risky actions wait for a person. When the signed log proves a loss, Breslau pays it up front and chases the money afterwards. Open the live desk See the rates Payables agent policy BRS-28-93858 Simulated Premium accrued today $213.4480 Actions 2,006 Held for review 2 Seq Tool Value Decision Entry hash Ed25519 signed, SHA-256 chained chain intact 400 tests pass on the meter software 1,000,000 log entries built and verified in the benchmark 3.5 bp of a payment's value is the illustrative premium 0 claims the software can refuse without a person A thought experiment An insurer designed from a blank sheet for AI agents Breslau started as a thought experiment. Design an insurance company from scratch, with AI doing the work people do today and a business model unlike the one insurers run now. To keep it concrete, the exercise picked a single market: US companies whose AI agents move money. That market is opening now. Agents can already pay invoices and place orders, and the forms behind standard liability policies now let carriers write generative AI out of cover. Somebody will have to price the risk. Breslau prices every action as it happens and pays claims from a signed record, with people only where the law requires them. The brief, and what the design did with it The brief asked for The design Built from the ground up A new line of cover for AI agents that move money A new business model Premium charged per action, plus a fee for the meter, much like a card network No humans Nine agents run the work. US law still needs five accountable people. Very profitable A 41% EBITDA margin in the 2031 base case, on assumptions no one can test yet A design on paper hides its weak points, so the parts that could be built were built. The meter runs as working software with 400 passing tests. Behind the numbers on this page sit a pricing model of 40,000 simulated years and a five-year financial model. Some of it does not hold yet. Nobody has measured how often an agent pays the wrong party, so the rates rest on judgment, and US law puts licensed people at points the design would rather automate. Those limits appear on the page wherever they bite. The gap Liability insurers can now exclude generative AI A company whose carrier adopts the new wording carries its agents' mistakes on its own balance sheet. ISO, the Verisk unit that drafts standard policy forms, has released endorsements that remove generative AI from commercial general liability cover. Carriers are filing to use such wording with state regulators, according to Insurance Journal . Specialist insurers already sell AI liability cover, priced for the agent as a whole from audits and scheduled tests. Breslau prices each action as the agent takes it, using what the meter records. How it works A meter in the path keeps a signed record of every call Customers install the meter and nothing else. Breslau bills premium and settles claims from the log it writes. Your agent calls a tool The meter prices each call and can hold it Tools and bank money and records move Premium billed monthly from each action's charge Advance paid once the log proves the loss Recovery from whoever ended up with the money The meter It runs as a gateway between the agent and its tools, or reads from a gateway the company already has. Every entry is chained to the one before and signed. Every ten days the meter checks its log against the bank statement. The price Each action adds to the premium. A payment costs basis points of its value and a commitment . The meter itself carries a separate fee of basis point. The cover Some losses sit in the log beyond argument, such as a payment made twice or one sent to a party the customer never approved. Breslau would pay those within hours, a design target, then recover the money from whoever received it. The brake The meter stops a payment above the agent's limit, a payee nobody approved, a repeat, or a payment split to slip under the limit. Someone at the customer then lets it through or kills it. Policies that switch the brake off pay times the rate. Rates Illustrative rates, shown in full A basis point is one hundredth of one percent. An agent's examined score moves every rate between 0.7 and 2.2 times base. Rate card rating plan Action What it covers Base rate Pay Moving money bp of value Commit Placing orders and signing contracts bp of value Message Messages to people outside the company $ per 1,000 Write Changing a record $ per 1,000 Read Retrieving records $ per 1,000 Authorization fee Charged on payments and commitments to run the meter bp of value Insured premium Metering charge with no cover attached Minimums $ premium and $ in fees a month Payments your agents make in a year $10M $500M $5B The brake On Off Premium and fee, a year Typical yearly loss to payment errors at 0.1% of spend The error figure is the low end of the 0.1 to 0.5% of spend that one recovery audit firm reports for duplicates and overpayments in human-run payables. Nobody has published the figure for agents. Where the plan rate sits against expected loss Low to high scenario Central Plan rate Shown in basis points on a log scale. Nobody publishes how often an AI agent pays the wrong party, so the scenarios start from data on human payment errors and fraud, and every step after that is judgment. With the brake on, the low and high scenarios sit a hundredfold apart. Risk What the pricing model expects to pay A simulation of 1,000 policies over 40,000 years, run at the plan rates. Hover any mark for its value. The brake and recovery remove most of the loss Payments, central scenario, basis points of value Grey bars show the loss before any control. The colored bars show what Breslau still pays after recovering what it can, first with the brake off and then with it on. Loss ratio across 40,000 simulated years Base book, 70% of policies on the brake The loss ratio averages and passes 100% in about one year in . Most of the long tail comes from not knowing the true error rates, which only a year of metering can narrow. What moves the price most Indicated payment rate, brake on, bp at a 55% loss ratio Low-risk setting High-risk setting The price is most sensitive to how often an agent pays the wrong party, a number nobody has measured yet. Economics A small company on a large flow Figures come from the financial model's base case, built from the rating plan and the assumptions written out in the workbook. Revenue, 2031 $37.4M on $250 billion of payments EBITDA, 2031 $15.3M 41% of revenue Peak funding $24.7M start of 2031 Break-even flow $117B a year of authorized payments Revenue by line, and EBITDA Base case, millions of dollars Meter fees start in 2027. Insurance income follows in 2028, and Breslau's share of the risk climbs from 5% of premium to 20% by 2031. EBITDA in 2031 by payment flow Costs stay almost flat as volume grows, so each extra billion of payments lifts the margin. Cumulative EBITDA by case The downside case runs at 30% of base flow with a 75% loss ratio and is still losing money in 2031. Software The meter already runs and has been stress-tested Written in plain JavaScript with no outside packages. In its end-to-end demo the meter drives 600 tool calls through thirty simulated days with the brake on, then a second short run with it off. The agent repeated an $8,400 supplier payment Held Flagged as a duplicate. The customer rejected it and no money moved. An injected instruction asked for $240,000 to an unknown company Held It broke both the limit and the payee list. $18,500 went straight to the bank, around the meter Uninsured The bank reconciliation caught it and the premium was corrected. With the brake off, a $6,250 invoice was paid twice Advanced Two log entries proved it, and the claims code advanced $5,250 without a person. Someone changed one digit in the log afterwards Detected Verification failed at that exact entry. Every entry carries the fingerprint of the one before Illustration of four log entries Entry 604 bank.pay $6,250.00 prev 9f2c41d0 hash 3a7be915 Signed Entry 605 mail.send n/a prev 3a7be915 hash c41e0b72 Signed Entry 606 bank.pay $6,280.00 prev c41e0b72 hash mismatch Fails here Entry 607 erp.read n/a prev e80d5a3f hash 71b2c6aa Signed Each hash covers the entry and the hash before it, and each entry is signed. Change the $6,250 payment to $6,280 and verification stops at that entry. A stress test then went after the software. It found three request shapes that reached the tool with no log entry and a claim path that paid on an unsigned log. All four are fixed and covered by tests. The live desk runs the same library in your browser, and its verify view checks a log the command-line meter signed. The meter runs on the customer's machine, so a customer can withhold a whole log. No outside security firm has reviewed the code. Plan A year of metering comes before the first policy A rate that could be wrong a hundredfold argues for counting first. The design sells the meter alone for a year and insures nothing, so that a loss table exists before a carrier is asked to trust a price. 2027 Meter only Brake and recovery service at design partners. The customer carries its losses, as today. 2028 Insurance Cover issued on a surplus lines carrier's paper, with reinsurers behind it and Breslau holding a small share. After 2031 A pool the customers own A reciprocal exchange owned by the insured companies and managed by Breslau for a fee. Share of flow that is insured Breslau's share of the risk Why the name The meter writes the table of agent failures that is missing today. In 1693 Edmond Halley took five years of birth and death registers from the city of Breslau and produced the first usable life table. Annuities could be priced properly from then on. The opening rows of Halley's Breslau table, from An Estimate of the Degrees of the Mortality of Mankind, 1693. Age Persons In 1835 Zachariah Allen fitted his Rhode Island mill against fire and asked his insurer for a lower rate. The insurer refused, so Allen formed a mutual with other owners who had made the same improvements. Breslau's third stage follows his model. Company Nine agents and five accountable employees Agents do all the routine work in the design, each inside a written limit. US insurance law requires accountable people. The design starts with four and adds a licensed producer before the first policy. A contracted licensed adjuster decides every claim the Adjuster agent would refuse, since its code has no deny outcome. Chief executive From 2027 Finance officer From 2027 Engineer From 2027 Engineer From 2027 Licensed producer Before the first policy Each agent has a written charter. Only the Adjuster's limit is enforced in code today, and none of the charters has run. How work moves between the agents Each box shows the limit written into that agent's charter Breslau is a design with software that already runs Nobody has incorporated or licensed Breslau. The live desk runs an example agent for a month on the real engine. You can break it five ways, or send the whole book through 20,000 simulated years. Open the live desk Breslau Breslau is a working name for a concept. It has no incorporation and no authority to transact insurance anywhere. Nothing on this page offers or solicits insurance. All agents, rates, figures and meter readings shown are illustrative, and AI agents carried out the reviews and tests described. Sources: Insurance Journal, 17 August 2026, on ISO endorsements CG 40 47, CG 40 48 and CG 35 08. Halley's table appeared in the Philosophical Transactions of the Royal Society in 1693. Zachariah Allen founded the Manufacturers Mutual Fire Insurance Company in 1835. A side project by Dinand Tinholt . More at dinand.com/side-projects . ### Chicago Downtown Walk URL: https://dinand.com/side-projects/chicago-downtown-walk/ Section: Side Projects · 3D city · OpenStreetMap data A walkable 3D model of downtown Chicago. Slide the light from dawn to midnight, or switch on the rain and watch the streets go wet. Best on a laptop. Drag to look around, use W A S D to walk and F to fly.