Robot evaluation methodology

Trust what you can inspect.

A benchmark earns credibility through its task design, execution, and evidence. Our approach makes those choices visible, from the industry workflow to the final trial record.

Start with a workflow, then bound the claim

A useful task begins with a clear account of the work: the objects, environment, handoffs, and outcome that matter to the people performing it. Industry conversations inform the benchmark design. They do not, by themselves, validate a benchmark or establish an industry-wide standard.

Separate the part a robot can be evaluated on from the surrounding workflow. Record what the task represents, what was simplified for repeatability, and what remains outside the test. Success on that task is evidence about the task—not a claim that the entire workflow is automated.

Make the embodiment gap visible

There can be three different configurations: the robot used to develop the model, the robot used for evaluation, and the manufacturer’s intended robot. Sharing an embodiment category does not make them interchangeable.

Document the robot model, end effector, sensing, calibration, controller, action space, control rate, compute, and software versions. Establish whether each policy can run on that setup. Record any adapter, fine-tuning, or other preparation separately from scored trials.

An integration failure is different from a task failure. If a result is meant to inform an OEM integration, identify the remaining configuration differences and define what must be tested on the target system. A single score should never hide that gap.

Freeze the protocol before scoring

Define the task, initial-state ranges, reset tolerances, observable success criteria, timeouts, and rules for human assistance. Specify the policies, conditions, trial allocation, order, exclusion rules, and reporting plan before the scored block begins.

Use a pilot to check feasibility, calibration, recording, scoring, and cycle time. Keep pilot and tuning runs separate from the evaluation. When the protocol changes, assign a new version and explain the effect on comparability.

Account for every scheduled trial

Keep a ledger of scheduled, attempted, completed, interrupted, excluded, and repeated trials. Assign stable identifiers that connect recordings, execution logs, outcomes, and interventions. A rerun does not erase the original attempt.

Apply prespecified rules consistently and preserve the reason for an exclusion. Record autonomous and assisted completion separately. A hardware interruption remains in the ledger even when the scoring rule excludes it from a particular rate.

Report the denominator and the uncertainty

Show results by task and condition, including sample counts and the definition of each metric. Report timeouts and failures alongside timing summaries. Describe uncertainty using a method appropriate to the design; trials within the same session or workcell may be related.

Keep observed outcomes distinct from suspected causes. An aggregate improvement can conceal a task-specific regression. Consequential or ambiguous findings may need a separately planned confirmation session rather than additional trials chosen after seeing the result.

Make public findings inspectable

For the first public benchmark, our intended release standard is an open protocol and results, a complete trial accounting, and supporting evidence that lets others examine the conclusions. Any limits on access to recordings, checkpoints, or logs should be stated alongside the result.

Agree on sharing and publication terms before testing. Disclose relevant funding, paid evaluation relationships, preparation budgets, and deviations from the protocol. Preserve unfavorable and inconclusive findings with the same care as favorable ones.

Methods evolve. A public release should identify its version, corrections, and limitations so later work can build on the same record. This page describes our approach; it is not a report of completed benchmark results.

Related research

Robot evaluation is an active research area. These projects offer useful perspectives on physical comparisons and the relationship between simulated and real-world evidence.

These are references, not Coop partnerships or endorsements.

Put the question first

What does your next model need to demonstrate?