For robot model developers
Robot regression testing across model versions
What improved in the new model, and what stopped working?

Keep a reference as the model changes
Model development produces a sequence of checkpoints, data changes, and control updates. A regression evaluation asks whether a candidate preserves previously demonstrated capabilities while improving the tasks it was intended to address.
Coop’s approach uses a versioned task suite drawn from the workflows that matter to the developer. A reference policy and a candidate are evaluated under the same scoring rules, with the software and physical configuration recorded for each run.
Keep the comparison stable
We document calibration, reset tolerances, task distribution, and trial scheduling. Where feasible, reference and candidate trials are interleaved within the same sessions so a historical baseline is not compared blindly with a changed workcell.
Changing the robot, camera setup, scoring rule, or task difficulty can break comparability. Those changes receive their own protocol version. Adaptation and tuning budgets are recorded separately so improvement is not attributed to a checkpoint alone when preparation also changed.
Look beneath the aggregate score
A higher overall success rate can conceal a lost capability on a particular task. The comparison therefore reports per-task and per-condition outcomes, autonomous and assisted completion, and the evidence behind any recurring regression.
A small apparent change may be inconclusive with the available trials. The evaluation plan should define the decision and confirmation procedure before results are examined. Repeatedly adding trials until a preferred version wins does not produce a reliable release decision.
- Reference and candidate checkpoints identified explicitly.
- A stable task suite and versioned physical setup.
- Per-task changes with counts and uncertainty.
- Follow-up tests defined for consequential or ambiguous findings.
Questions about regression testing
Can this support repeated model releases?
Yes. A versioned suite provides a reference for later comparisons. The execution schedule, compatibility, and scope of each evaluation are agreed with the model developer.
Can we reuse an old baseline?
Only with its limitations made explicit. A changed physical setup or session can affect results. Re-running the reference alongside the candidate usually provides a more interpretable comparison when feasible.
Does Coop set our release threshold?
The decision criteria are agreed before testing. A report provides scoped evidence and uncertainty; the developer remains responsible for the broader release and deployment decision.
Further reading
Measuring robot task success and human interventions