A benchmark becomes useful to an industry when its tasks reflect work that matters. That connection needs to be designed and documented. An industry label attached to a familiar manipulation task is not enough.

The starting point is a conversation with people who know the workflow. The output is a bounded evaluation question for a model developer, with the assumptions and simplifications made visible.

Understand the outcome and the surrounding work

Ask what needs to happen, what counts as an acceptable result, what changes between instances, and what people do when a step fails. Understand the inputs, handoffs, timing expectations, and operating conditions before selecting a robot task.

Record whose perspective informed the task and which parts of the workflow remain uncertain. A conversation is an input to design, not a claim of partnership or approval from an entire sector.

Choose a task with an observable boundary

Consider a hypothetical workflow that involves selecting a requested item and placing it in a designated tray. A bounded benchmark could begin with a documented arrangement of items and end when the selected item reaches a defined region, or when a timeout or failure condition occurs.

That task would need explicit rules for item identity, acceptable placement, distractors, retries, and human assistance. It would not establish performance on the upstream inventory process or downstream handoff. The example is a design illustration, not a claim of a completed Coop benchmark.

Represent relevant variation deliberately

Choose variations because they matter to the workflow: object arrangements, the presence of distractors, or a defined change in lighting. Document the reference condition, the range being tested, and what remains fixed.

One-factor tests can make a failure easier to interpret. Combinations answer a different question and need their own trial allocation. Avoid presenting a convenient set of lab variations as a representative distribution of real operating conditions unless there is evidence for that distribution.

Choose hardware without erasing the embodiment gap

A task can be meaningful while a particular model is incompatible with the proposed robot. Qualify the policy interface and physical setup before claiming a comparison is possible. End effectors, sensing, calibration, and control characteristics can change the task itself.

If the intended manufacturer uses a different configuration, document that difference. Cross-embodiment evaluation can be valuable, but it does not isolate the effect of a policy in the same way as comparing compatible models on one fixed setup.

Publish the rationale beside the protocol

A benchmark release should explain why its tasks were chosen, how they relate to the source workflow, and what was simplified. Pair that rationale with reset rules, scoring definitions, configuration details, and a version history.

Invite corrections to the design and distinguish them from changes that would make results incomparable. A better task definition may justify a new benchmark version; it should not silently change the meaning of an existing score.

Put the method to work

Coop

Independent robot benchmarks, grounded in real work