Working research productSimulation with a browser walkthrough

Physical AI and oversight

Pakhuis

A seventeenth-century Amsterdam canal warehouse, built in simulation, where two robots count stock and a reviewer decides what reaches the record.

The question behind the system

When a human reviews a robot's work before it counts, how far does the robot roll and how long does it stand still while it waits?

What it is designed to do

  • Price human review of robot work in seconds and metres, with every posted number scored against simulator ground truth
  • Check a vision model's stated confidence against the objects it recognises
Campaign
3 policies × 3 reviewer rates, registered before launch
Scope
Two Nova Carter robots, one deck, 11 or 12 counted locations
Per run
Every count, verdict and halt event logged
Bench
Oversight bench, 9 registered cells

What it found

With every count reviewed and the robots still driving, review caught the one fault the unreviewed robots had posted as correct, so all five planted faults were caught.

Stopping a robot for its verdict took about 0.2 seconds and a quarter of a metre at every reviewer rate. The wait that followed averaged 6.8 seconds with the slowest reviewer and 1.3 with the fastest.

The robot's vision model reported 81 to 85 percent confidence on frames where its recall ran from 37 to 64 percent, in all eight test conditions.

Two of the five registered predictions, yield (H1) and throughput (H3), failed as worded; wrong posts (H2) and pending posts (H5) held, and halt latency (H4) held except its order-of-magnitude clause at the fastest rate. Halting changed which bays the robots could read, and at the fastest reviewer rate the halting fleet patrolled faster than the unreviewed one, for reasons the run files do not show.

The ground-floor aisle of the Pakhuis, path-traced, with crates in the foreground and barrels at the rear wall.
Ground floor, camera A. Path-traced still from the simulation.
The project film, 86 seconds.

The public boundary

What this page does—and does not—open.

The walkthrough is a self-contained page of about a megabyte that runs in the browser with no server, while the simulation itself stays on a dedicated workstation. The reviewer in this first test was scripted to be always right, and each condition ran once, all of it in simulation.

Walk the building in your browser See the Assay Bench profile
All experiments Back home
“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top