Review standard · Version 1.1

Same questions.
Visible evidence.

A consistent rubric makes the review useful to readers and predictable for manufacturers. Every score must point to a recorded observation.

Three categories. 100% of the score.

What the rating measures.

Judged against the product’s stated purpose and intended user. A learning robot and a household robot face different jobs.

01

Quality & Value

Is it well made—and worth the total cost? Inspect construction and consistency, then judge the tested capability against the price, required extras and realistic alternatives for its intended user.

Evidence & scoring anchors

Hardware inspection, repeated use, observed wear and faults; dated price, required accessories and task-matched alternatives.

  1. 0 / 5Hardware faults or the tested capability prevent the product from serving its stated purpose at the required cost.
  2. 1 / 5Frequent hardware problems or a poor match between capability and total cost leave major barriers for the intended user.
  3. 2 / 5The product serves a narrow use, but observed construction issues, inconsistency, required extras or cost materially weaken its value.
  4. 3 / 5Construction and tested capability reasonably suit the intended purpose and total cost, with clear compromises.
  5. 4 / 5Repeated use shows sound construction and consistent operation; dated comparisons support strong value for the intended user.
  6. 5 / 5Documented use supports exceptional quality and a compelling task-matched value after accounting for cost, extras and remaining durability uncertainty.
40%of overall rating
02

Usability & Customization

Can you use it reliably—and make it your own? Follow setup, everyday operation, recovery and useful modifications. Repeat tasks with changed conditions to show where behavior generalizes, where it fails and how much human help it needs.

Evidence & scoring anchors

Setup time and prerequisites; baseline and held-out runs, failures, resets and interventions; a documented modification and restore procedure.

  1. 0 / 5The intended user cannot reach a working core task through the supported setup and control path.
  2. 1 / 5Use depends on expert workarounds or a carefully tuned scene; ordinary changes break behavior and meaningful customization is blocked.
  3. 2 / 5Baseline tasks work, but setup, recovery, changed conditions or modifications require substantial effort or extra assistance.
  4. 3 / 5Setup, repeat use and a purpose-relevant modification are reproducible for the intended audience; tested variations establish clear limits.
  5. 4 / 5Clear instructions and recovery support straightforward use and customization, with consistent results across meaningful tested variations.
  6. 5 / 5A fresh-start walkthrough reproduces setup, extensions and restoration; independent repeats support a broad tested envelope in the declared control mode.
40%of overall rating
03

Support & Ownership

What is it like to keep, maintain and get help with? Evaluate documentation, replacement parts, maintenance, updates and actual support interactions. Make ongoing costs, restrictions and dependencies clear.

Evidence & scoring anchors

Dated documentation and support records; practical parts, repair and update paths; ongoing costs and restrictions.

  1. 0 / 5An observed essential ownership problem has no workable documented or tested support, repair or replacement path.
  2. 1 / 5Essential documentation, parts or recovery support is demonstrably missing or ineffective, leaving an unresolved barrier to continued use.
  3. 2 / 5Resources are available, but documented gaps, restrictions or unresolved support interactions require significant owner workarounds.
  4. 3 / 5Documentation and practical maintenance, parts and update paths support ordinary ownership; known costs and limits are clear.
  5. 4 / 5Dated evidence shows useful support and workable repair or update procedures, with accessible resources and transparent ongoing obligations.
  6. 5 / 5Documented issue resolution and ownership procedures establish dependable support, parts access and recovery across the needs actually evaluated.
20%of overall rating
The score is a summary, not the evidence.

A published 0–5 scale.

0Unable to meet the defined task; demonstrated failure.
1Major limitations; frequent help or unreliable outcomes.
2Works narrowly, with substantial constraints or effort.
3Meets its purpose under documented conditions.
4Strong results across meaningful variations.
5Exceptional for its purpose, with broad supporting evidence.
Untested is unrated. We publish an overall score only when all three categories have supporting evidence. We never replace missing evidence with zero or silently redistribute weights.

Overall = Σ (category score × weight) ÷ 100, rounded to one decimal. A serious safety concern or unresolved essential failure is called out separately and can prevent a recommendation regardless of the average. We do not issue safety certification.

Explore the formula

Try the rating calculator.

This example explains the weights. It does not assign a rating to any robot.

Illustrative calculation only
Unrated

Select all three categories to see the weighted example.

The generalization check

Beyond one good take.

A robot can repeat a rehearsed sequence and still struggle with a small change. The review needs to show where that boundary lies.

A declared task, then a frozen 20-run plan

Where feasible, select a purpose-relevant core task and plan 20 measured attempts: 10 baseline and 10 held-out attempts spanning at least three declared variations.

Twenty planned attempts are a practical review protocol, not a statistical generalization claim. Conclusions apply to the tested task, conditions, control mode and actual sample. Score each dimension from its observed evidence; missing or inapplicable evidence remains unscored, never zero.

Make the test reproducible

Record the exact hardware, software, environment and task. Set pass criteria before scored runs. Report the actual attempt count and successes, rather than selecting the best clip.

Change something meaningful

Use held-out positions, objects or lighting where the task permits. Keep development and measured trials separate. Explain the range tested—and avoid claiming general-purpose behavior from a few variations.

Account for human help

Label teleoperation, scripted execution, replay and learned policy inference. Report resets, recovery work and interventions. Partial success and assisted success must stay visible.

Keep the evidence current

Date the verdict and software configuration. Updates that change behavior require new tests. A brief review cannot establish years of durability.

Read the editorial policy