For robot model developers

Robot policy evaluation, grounded in real work

What can your model reliably do on the tasks that matter to a workflow?

Our methodology
Concept illustration of matching humanoid workcells used for a controlled policy comparison
Concept illustration. Hardware is scoped for each evaluation.

A useful comparison starts with a useful task

Model developers need evidence that connects a policy’s capabilities to work someone actually wants a robot to do. Coop’s approach starts with conversations about those workflows, then turns a bounded part of the work into a task with observable success criteria.

The evaluation can characterize one runnable model or compare compatible policies on the same robot configuration. A common protocol makes the result useful for model development and for technical conversations with robot manufacturers considering a model integration.

Compare the complete deployed policy

Before scored trials, we establish the robot, end effector, cameras, controller, compute, and policy interface. A checkpoint alone does not define what is being compared: preprocessing, action scaling, control frequency, and any adaptation are part of the tested configuration.

Policies share the task definitions, initial-state ranges, timeout rules, and assistance policy. Trial order is balanced where appropriate so a later session, drifting calibration, or a changing workcell does not quietly favor one model. Integration and tuning runs are recorded separately from scored evaluation.

Measure outcomes that can be inspected

Autonomous task completion, completion time, and human interventions answer different questions. We report them separately, by task and condition, with trial counts and uncertainty. Failures and timeouts remain visible rather than disappearing behind an average of successful runs.

The report connects each conclusion to the relevant trial records and recordings. When a difference is too small or inconsistent to support a claim, that uncertainty is part of the result.

  • Success under a prespecified, observable rule.
  • Assisted completion and interventions recorded separately.
  • Failure patterns connected to individual trials.
  • Scope and limitations for the tested embodiment.

Questions about policy evaluation

Can we evaluate a single model?

Yes. A baseline evaluation can establish what a runnable model does on the chosen tasks. Comparing it with another model requires both policies to be compatible with the same documented setup.

Does a result establish performance on an OEM’s robot?

It establishes performance on the tested configuration. A manufacturer’s robot can differ in its gripper, sensing, controller, or other properties, even within the same embodiment category. Those differences need a separately scoped evaluation.

Which robots and models are supported?

Compatibility, hardware availability, integration work, timing, and deliverables are established for each evaluation. A robot brand or model checkpoint alone is not enough to confirm support.

Further reading

How to evaluate a robot policy on real hardware