← Back to Countback

MEASURED ALGORITHM BEHAVIOR

Fewer additional counts in 662 of 1,944 generated cases.

We compare each policy’s realized path on the same generated truth. The opening, transfer model and solver stay fixed. One arm also has the truthful historical observation.

Paired outcomeCasesShare
Fewer additional counts with history66234.1%
Same additional counts1,28065.8%
One extra count with history20.1%

All 240 intentionally unsupported controls refused or reported a contradiction.

What this does not prove

These are synthetic cases, not customers. Historical observation acquisition, setup, movement verification, walking time and review time are excluded. Some generated observations occur at the final boundary. The benchmark does not establish net labor savings, warehouse ROI or product-market fit.

An earlier report’s 57.1% figure describes a lower lexicographic planner objective. Its summed-path component compares different feasible-state sets; it must not be read as 57.1% of cases saving physical work. The paired figures above are the user-facing claim.

Reproduce and inspect

Derived summary and source SHA-256 · Complete original per-case results · Original historical report

python -m scripts.build_judge_site

The builder derives the paired summary from every row in the original results; seed 20260906. Source and benchmark generation scripts.