Every humanoid robot announcement comes with footage. The robot walks across a room, picks up a box, hands a tool to a waiting worker. The clip is edited cleanly, the movement is smooth, and the implied message is clear: this system is capable, maturing, nearly ready. What the footage almost never shows is how many attempts it took, what percentage of tries succeeded, what happened when they didn't, and by what standard anyone decided the robot was performing well enough to show publicly.
That missing context matters more than most coverage acknowledges. The question of how you evaluate a humanoid robot — not in a demo reel, but rigorously, against consistent criteria — is one of the least-discussed and most consequential open problems in the field. Without agreed benchmarks, "deployment-ready" can mean almost anything a company wants it to mean. And right now, there are no agreed benchmarks.
Why Evaluation Is Harder Than It Looks
Testing a humanoid robot is complicated by the same property that makes humanoids attractive in the first place: they are supposed to be general-purpose. A robot designed to do one thing — sort packages, weld a specific joint, move a specific type of bin — can be evaluated against a precise, narrow specification. Does it complete the task? At what speed? With what error rate? How often does it require human intervention?
A humanoid robot is meant to operate across many tasks in environments that vary. Evaluating it comprehensively requires a test suite that covers the range of things the robot is supposed to do, under the range of conditions it will actually face. That is both technically demanding and commercially sensitive — companies testing their own robots have strong incentives to design tests that their robots perform well on, and to publish results selectively.
There is also a fundamental gap between laboratory performance and field performance that any honest evaluation framework has to grapple with. A robot can achieve high task success rates in a controlled test environment where every variable is managed, and fail at much higher rates when deployed in a real facility with unexpected obstacles, variable lighting, and the unpredictability introduced by human co-workers. The history of industrial automation includes many systems that passed acceptance testing comfortably and then underperformed in production. Humanoid robotics is unlikely to be an exception.
The Metrics That Matter — and the Ones That Get Used
The metrics most commonly cited in humanoid robotics coverage are not the ones that matter most for deployment decisions.
Speed of movement and agility in demonstrations — how fast the robot walks, how fluidly it moves — gets disproportionate attention because it is visually striking and easy to film. It is not, for most commercial use cases, the variable that determines whether deployment makes sense. A warehouse robot that moves at 70% of human walking speed but maintains high task completion rates is more commercially useful than one that moves faster but fails one task in five.
Task success rate — what fraction of attempts at a defined task the robot completes correctly without human intervention — is the metric that matters most for deployment economics, and it is the one companies are least willing to publish. When success rates do appear in company materials, they are almost always measured under controlled conditions on a task the system has been specifically trained and tuned for. Generalisation performance, measured on tasks or variations the system has not been specifically prepared for, is rarely reported.
Mean time between failures, and mean time to recover from a failure, matter enormously for operational planning. If a robot fails at a task and takes two minutes of human intervention to reset and resume, that has very different operational implications than a failure that takes thirty seconds or one that requires a maintenance technician. These figures are essentially never reported publicly.
Uptime — the fraction of scheduled operating time the robot is actually functioning — is another critical operational metric that companies consistently omit. A robot that performs well when it is running but requires frequent maintenance stops is not deployable at scale in the way that a machine with 95% or higher uptime might be.
Academic Benchmarks and Their Limits
The robotics research community has developed a range of benchmark tasks and datasets intended to provide consistent evaluation criteria. The most widely used for manipulation — the set of tasks involving picking up, moving, and placing objects — include tasks drawn from everyday scenarios: picking objects from a bin, setting a table, folding laundry. Research groups publish performance figures on these tasks, and the results allow some comparison across systems.
The limitation of academic benchmarks is that they are optimised for comparability across research systems, not for predicting commercial performance. A benchmark task designed to be tractable enough that multiple research groups can attempt it is, by construction, a task that is easier and more controlled than the conditions of an actual deployment. Progress on benchmark tasks is real — performance on grasping benchmarks has improved substantially over the past five years — but the relationship between benchmark performance and what a robot can reliably do in a messy real-world environment is loose.
Some researchers have argued for what they call "real-world benchmarks" — evaluation frameworks that test robots in uncontrolled environments, with real variability, rather than in carefully constructed laboratory settings. These are methodologically harder to run and harder to compare across institutions, but they provide more honest signal about deployment readiness. They remain a minority of published evaluation work.
What Operators Actually Use to Make Decisions
Companies that have deployed or piloted humanoid robots — including the handful that have announced real operational programmes — are generally not using published academic benchmarks to make deployment decisions. They are running their own internal evaluations, designed around the specific tasks and environments that matter for their operations.
In practice, this means that the most meaningful evaluation work happening in the industry right now is not publicly visible. Amazon's evaluation of Agility Robotics' Digit, for example, almost certainly involved proprietary testing against specific warehouse task specifications, uptime requirements, and safety criteria. None of that is published. The deployment decision reflects a private judgement that the system meets a private standard.
This is not unusual in industrial equipment procurement — companies evaluating any piece of capital equipment run their own tests rather than relying solely on manufacturer specifications. But it does mean that the public record on humanoid robot performance is substantially thinner than the volume of coverage suggests. There is a lot of footage and a lot of press releases. There is very little independently verifiable performance data.
Safety Evaluation Is a Separate Problem
Performance evaluation — can the robot do the task? — is distinct from safety evaluation — can the robot operate without injuring the humans working near it? Both matter for deployment, but they require different frameworks and are at very different stages of maturity.
Industrial robots in fixed installations have well-established safety standards, developed over decades by standards bodies including ISO and the American National Standards Institute. These standards define requirements for guarding, emergency stop systems, safety-rated control systems, and testing procedures. They work because conventional industrial robots operate in defined, enclosed workspaces with predictable trajectories. A robot arm that never leaves its programmed range of motion is much easier to make safe around humans than one that moves through a shared space unpredictably.
Humanoid robots operating alongside humans in shared spaces — which is the core use case — require a different approach to safety. The relevant standards are still being developed. ISO Technical Committee 299, which covers service robotics, has been working on standards for mobile robots operating in human environments, but these are not yet mature for humanoid-specific deployment scenarios. Several humanoid companies have their own internal safety certification processes, but these are not standardised or independently audited in the way that established industrial safety standards are.
This is not a criticism unique to humanoid robotics — it reflects the general state of standards development when a technology moves faster than the standards bodies tracking it. But it is worth being specific: when a company says its humanoid is "safety-certified," that claim requires scrutiny about what standard was applied, by whom, and under what conditions.
What a More Honest Evaluation Framework Would Look Like
Several researchers and industry observers have called for the development of agreed industry benchmarks — something analogous to what ImageNet provided for computer vision in the early 2010s: a common test suite that allows meaningful comparison across systems and tracks genuine progress over time.
For humanoid robotics, the practical shape of such a framework would likely include standardised task sets covering manipulation, navigation, and interaction with human co-workers; evaluation under controlled and uncontrolled conditions; independent evaluation rather than self-reporting; and metrics that go beyond task success rate to include failure modes, recovery behaviour, and operational uptime.
The obstacles are real. Companies have strong competitive incentives not to expose their systems to independent testing, particularly in areas of weakness. Building test environments that are representative of actual deployment conditions without being so uncontrolled as to be unreproducible is technically hard. And the humanoid robot space is still early enough that the population of systems capable of attempting a meaningful benchmark task suite is small.
Still, the absence of agreed benchmarks carries costs that are easy to underestimate. It makes it harder for potential operators to make informed procurement decisions. It makes it easier for companies to control narratives around their systems' capabilities. And it makes it genuinely difficult to assess whether the field is advancing as fast as the announcements suggest.
The companies that will build lasting credibility in this space are the ones that publish real performance data — task success rates across varied conditions, uptime figures from actual deployments, honest accounts of failure modes — rather than curated footage of their best attempts. That hasn't happened in a meaningful way yet. When it does, it will be more informative than any demo video. The industry's willingness to be evaluated honestly is itself a signal worth watching.