How to Define Success for an Industrial Humanoid Pilot
Without written acceptance criteria, a pilot can remain an impressive demonstration with no production decision.
A pilot should end with a decision. That requires the evidence for success, redesign and stopping to be agreed before the team begins engineering.
Separate demonstration from production evidence
A demonstration proves that a behavior can happen under selected conditions. A production pilot tests whether the task can be performed repeatedly inside a defined operating envelope with acceptable intervention, recovery, safety and support.
Both are useful, but they answer different questions.
Define the operating envelope
Document the parts, routes, machines, lighting, floor conditions, people, shift pattern and upstream/downstream assumptions included in the pilot. A result outside that envelope may be interesting, but it should not silently change the acceptance scope.
Use a balanced acceptance set
Task completion
Define precisely what counts as a successful cycle and how partial or recovered cycles are recorded.
Intervention frequency
Record human and remote-operator interventions by type. A task-completion rate can look strong while requiring economically unacceptable supervision.
Throughput
Use the operational unit that matters: cycle time, moves per hour, kits per shift or machines served.
Availability and recovery
Define planned operating time, excluded downtime, mean recovery time and what happens after lost localization, failed grasp or blocked path.
Variation coverage
List the product, container, machine, route and environmental variants that the evidence must cover.
Safety behavior
Confirm the intended operating modes, safe stops, restart process, access boundaries and residual risks. A pilot is not a substitute for the required formal safety and conformity work.
Sustained period
Set a minimum monitored operating period. Repeating a task for several shifts reveals different problems than a curated run.
Define decision thresholds
For each metric, agree:
- the target for continuing to scale;
- the range that triggers redesign and another test;
- the result that stops the application.
Do the same for economic and organizational assumptions. Include operator workload, maintenance ownership, support response and expected cost per successful task.
Written criteria reduce optimism bias and stakeholder disagreement. More importantly, they transform the pilot from a technology event into an industrial decision process.