
Why we’re building Coop
A successful demo is a starting point. We’re building the physical infrastructure to find out what works, how reliably, and under which conditions.
Read article
Coop evaluates robot policies on real hardware, measuring what works, where it fails, and how reliably it performs.


The physical side of AI
Understanding how a robot moves an object, recovers from a mistake, and works in changing conditions takes time with real hardware. Coop turns those questions into repeatable evaluations, with defined conditions, measurable outcomes, and evidence you can inspect.
Benchmarking & evaluations
Compare policies, test their limits, and measure progress through controlled experiments on real robots.

Evaluate compatible policies on the same robot, tasks, and starting conditions. Measure autonomous success, completion time, and interventions, with every trial accounted for.
How we benchmark
A useful benchmark makes clear what was tested, how it was scored, and what the result means. We build the evaluation around those details from the start.
Define the robot, tasks, conditions, and success rules. Check policy compatibility and establish a reference baseline before the scored evaluation begins.
Run repeated trials with controlled starting conditions and documented resets. Record failures, interruptions, and human interventions alongside successful runs.
Review results by task and condition, with trial counts, uncertainty, and recordings. Understand what the comparison supports, where it is limited, and what needs another test.
The Coop journal
The questions, methods, and physical work behind our benchmarks.