MEASURED ALGORITHM BEHAVIOR
Fewer additional counts in 662 of 1,944 generated cases.
We compare each policy’s realized path on the same generated truth. The opening, transfer model and solver stay fixed. One arm also has the truthful historical observation.
| Paired outcome | Cases | Share |
|---|---|---|
| Fewer additional counts with history | 662 | 34.1% |
| Same additional counts | 1,280 | 65.8% |
| One extra count with history | 2 | 0.1% |
All 240 intentionally unsupported controls refused or reported a contradiction.
What this does not prove
These are synthetic cases, not customers. Historical observation acquisition, setup, movement verification, walking time and review time are excluded. Some generated observations occur at the final boundary. The benchmark does not establish net labor savings, warehouse ROI or product-market fit.
An earlier report’s 57.1% figure describes a lower lexicographic planner objective. Its summed-path component compares different feasible-state sets; it must not be read as 57.1% of cases saving physical work. The paired figures above are the user-facing claim.
Reproduce and inspect
Derived summary and source SHA-256 · Complete original per-case results · Original historical report
python -m scripts.build_judge_site
The builder derives the paired summary from every row in the original results; seed 20260906. Source and benchmark generation scripts.