Open evaluation resources
Make the next
experiment inspectable.
Practical starting points for model developers and researchers planning physical robot evaluations. Download, adapt, and version them for your own protocol.
Evaluation protocol
Define the workflow, task, scoring rules, trial allocation, and publication plan before the scored run.
Download template JSON · v0.1Embodiment profile
Record the development, evaluation, and target robot configurations, including interfaces and remaining differences.
Download template JSON · v0.1Trial record
Connect an individual attempt to its protocol, conditions, outcome, interventions, and evidence files.
Download templateUse the three records together
- Write the protocol. Define the question and the decision the evaluation needs to support. Pilot the setup, then freeze and version the scored protocol.
- Identify the configuration. Complete a profile for the robot used in evaluation. Record the differences from the model’s development setup and any intended OEM configuration.
- Account for the trials. Copy the trial record for each scheduled attempt, using stable identifiers. Keep failures, interventions, exclusions, and reruns connected to the original record.
The JSON files are blank templates, not benchmark data. Replace placeholders with real values, keep missing measurements as null, and define any additional fields in your protocol. Arrays begin empty; the trial template includes a field guide for recording interventions.
Adapt the method to the question
These resources describe a starting structure, not a universal benchmark or a completed validation. Choose the trial allocation and uncertainty method for the actual design. A small set of repeatable tasks can answer a bounded question without establishing performance across an industry.
You may use, modify, and share these templates with attribution to Coop. The download files include the reuse terms. Feedback and corrections are welcome at founders@cooplabs.com.
Read the accompanying guides
- How to evaluate a robot policy on real hardware
- From an industry workflow to a repeatable robot benchmark
- Measuring robot task success and human interventions
