Industry-informed robotics benchmarks
Start with the work.
Make it measurable.
Independent benchmarks for humanoid and manipulation models, starting with warehouse workflows. We connect model performance to the tasks people need automated, with evidence developers and robot manufacturers can inspect.
Industry workflows
Talk with the people doing the work. Understand the task, operating conditions, and what a useful outcome means.
Repeatable tasks
Translate a bounded part of that workflow into a protocol with a compatible robot, reset rules, and observable scoring.
Model evidence
Run repeated physical trials and connect results to the conditions, failures, and interventions behind them.
Three evaluation questions
Understand a model. Then understand its limits.

Policy evaluation
What can your model reliably do on the tasks that matter to a workflow?
Explore the evaluation
Robustness testing
Which changes to a task expose the limits of your model?
Explore the evaluation
Regression testing
What improved in the new model, and what stopped working?
Explore the evaluationConcept illustrations of evaluation setups. Hardware and model compatibility are established for each scope.
Public benchmark development
Credibility starts with an inspectable record.
Our first public benchmark is in development. The release plan is to make the protocol, results, trial accounting, and supporting evidence open and verifiable, with any access limits stated explicitly.
These pages describe the approach and evaluation scopes. Public scores, model rankings, and completed industry benchmark results will appear only when the evidence is ready.
Read our methodologyEvidence for the next technical conversation
Our initial evaluation work is for robot model developers. Industry conversations help shape the tasks; the evaluations help developers understand their models and discuss capabilities with original equipment manufacturers (OEMs).
A result belongs to its tested embodiment. A similar-looking robot can have different end effectors, sensors, or control interfaces. Reports should make those boundaries clear before a result is used to support a model integration or licensing conversation.
