Robot Dexterity Testing: How the Benchmark Is Measured
Table of Contents

Robot Dexterity Testing: How the Benchmark Is Measured

Robot dexterity testing has four dimensions and most demos show just one. The benchmarks researchers use, and how to run a real one at your kitchen table.

Legs were robotics’ headline story for a decade. Hands are the story now, and the measurement problem arrived with them. When IEEE Spectrum covered Boston Dynamics’ new Atlas hand on October 1, 2026 (three fingers and a thumb, 13 degrees of freedom, up from 7), the article described the design reasoning in detail and specified no formal dexterity benchmark or testing protocol. That’s not an oversight by the reporter. It reflects a real state of affairs: robot dexterity testing exists, it’s well developed in research, and it is not yet how the industry talks about its own products.

Key Takeaways

  • A dexterity benchmark is a standardized set of physical objects plus a written protocol, so that two labs testing two robots are measuring the same thing.
  • The Yale-CMU-Berkeley object set (Calli et al., IEEE Robotics & Automation Magazine, 2015) is the reference example: everyday objects spanning “different shapes, sizes, textures, weight and rigidity,” distributed with RGBD scans, physical properties, geometric models and task protocols.
  • NIST runs the Robotic Grasping and Manipulation Competition and maintains task boards specifically to develop “performance metrics and test methods” for robotic assembly.
  • Dexterity has four separate dimensions: grasp variety, in-hand reorientation, force control, and durability. Demo videos almost always show the first and skip the other three.
  • The number that would settle most arguments is a success rate on a named object set. Research groups report it. Product announcements generally don’t.

What a benchmark actually is, and why physical ones are hard

A benchmark is a fixed test that lets you compare two systems by holding everything constant except the system. In software that’s easy: same input data, same metric, run it. In manipulation it’s brutally hard, because the test is physical objects in physical poses, and nobody has the same ones.

That’s exactly the problem Berk Calli, Aaron Walsman, Arjun Singh, Siddhartha Srinivasa, Pieter Abbeel and Aaron Dollar set out to fix with the YCB object set in 2015. Their approach was to pick a set of objects of daily life spanning different shapes, sizes, textures, weights and rigidities, then distribute both the physical objects and the digital assets (high-resolution RGBD scans, measured physical properties and geometric models), along with written protocols for evaluating planning, learning and control approaches.

Read what that solves. If my robot picks up a mug and yours picks up a mug, we’ve learned nothing. If both of us run the same protocol on the same YCB mug and report success rates, we can actually compare. The physical standardization is the contribution, not the mug.

NIST attacks the same problem from the manufacturing side. Its Robotic Grasping and Manipulation for Assembly project develops performance metrics and test methods for robotic assembly technologies, covering grasping, force control, assembly performance and dexterous manipulation, and maintains a practice task board plus a Manufacturing Track competition. The project’s framing is that “Recent advancements in robotic arms and end-effectors have the potential to accelerate the use of robotics for assembly”, and metrics are what turn potential into procurement decisions. NIST also calls out collaborative robots’ “force sensing and compliance capabilities used in collaborative robots to prevent injuries and enable them to work safely alongside human workers,” which links dexterity measurement directly to the safety methods we cover in what replaced the robot safety cage.

There’s a useful precedent outside robotics, too. Rehabilitation medicine has used timed object-transfer tests, which move standardized blocks or pegs between locations against a clock, for decades to assess human hand function. Those tests were built for people, not machines, and they illustrate the principle robots now need: same objects, same clock, reported number.

The four dimensions of dexterity, and which ones demos skip

Grasp variety. Can the hand hold many kinds of objects? This is what every video shows, and it’s the easiest dimension. It’s largely a function of finger geometry, friction material and grip force.

In-hand reorientation. Can the hand change the object’s orientation without putting it down? This is what separates a hand from a gripper, and it’s the capability that justifies the cost of fingers. Watch for it specifically: if the robot always sets an object down and re-grasps, it doesn’t have it.

Force control. Can the hand apply a specified force rather than just a position? This is what you need to turn a stiff knob without snapping it, to insert a connector until it clicks, or to hold an egg. It depends on sensing, and tactile sensing in robotics lags far behind vision because there’s no cheap, high-resolution, durable skin sensor.

Durability. Can the hand do all of that ten thousand times? This is the dimension the industry actually cares about and the one almost never shown. Boston Dynamics’ redesign is transparently about it: eliminate delicate tendons and cables, use fewer and larger direct-drive actuators embedded in the joints, make every actuator pack easily replaceable. Alberto Rodriguez, the company’s director of robot behavior, described the current work as figuring out what has to change “if we want to make 100,000 of these hands a year.”

Rodriguez’s framing of the whole exercise is the best one-line summary of the field: “Hands are a ruthless design trade-off…you’re always giving up on something.”

How to Teach Your Kid About Robot Dexterity Testing

This is the best hands-on benchmarking lesson available to a family, because the whole method fits on a kitchen table.

Ages 5–8: One hand, one rule

Have your kid pick up ten objects using only a thumb and one finger. Then only two fingers, no thumb. Then with a mitten on. Each restriction removes a capability, and they’ll notice which objects become impossible rather than just harder. Then the fun bit: time them. Timing turns “I can do it” into “I can do it in eleven seconds,” which is the entire idea of a benchmark.

Ages 9–12: Build the object set

Pick ten household objects deliberately spanning the YCB dimensions: shape, size, texture, weight and rigidity. A coin, a tennis ball, an egg, a pencil, a folded towel, a plastic bag, a wet glass, a key, a bag of rice, a sheet of paper. Write the list down, that list is now your family’s object set. Then write the protocol: five attempts per object, success means lifted 10 centimeters and held for three seconds, and record every attempt. They have just built a benchmark, which is a more sophisticated thing than most robot videos contain.

Ages 13+: Test the four dimensions separately

Using the same object set, run four tests with their own hand or a built gripper. One: how many objects can be grasped at all. Two: how many can be rotated 90 degrees without being set down. Three: how many can be held firmly without damage, the egg and the paper cup settle this. Four: pick the single hardest object and do it fifty times, counting failures. Then write up a one-page report with four numbers. Those four numbers are a more honest description of dexterity than anything in a product launch video.

The question to ask: “If a robot can pick up a hundred different objects but can never turn one around in its hand, how dexterous is it really?”

How dexterity gets measured, dimension by dimension

DimensionHow it’s measuredWhat a demo video usually showsWhat’s missing
Grasp varietySuccess rate across a standardized object setA montage of successful graspsThe denominator: how many attempts
In-hand reorientationCan the object be rotated without regrasping, and by how muchRarely shown at allWhether the hand can do it at all
Force controlTarget force versus achieved force; damage to fragile itemsRigid, forgiving objectsEggs, paper cups, wet glass, thin film
DurabilityCycles to failure, mean time between failuresNothingThe spec the buyer actually needs
GeneralizationSuccess on objects never seen in trainingObjects from the training distributionA declared held-out test set
SpeedCycle time per task, consistently measuredEdited or sped-up footageAn unedited continuous take
Safety of contactForces within biomechanical limits by body regionA person casually shaking the handWhich standard was applied

The pattern down that right column is the article. Almost everything missing from a robot hand demo is a denominator, a hard case, or a number.

What to do with this as a parent or a teacher

Ask for the denominator

“The robot picked up 50 objects” is not a result. “The robot attempted 100 grasps on the YCB set and succeeded on 73” is. Teaching a kid to ask “out of how many?” generalizes far beyond robotics, and it’s the fastest single upgrade to anybody’s media literacy.

Treat the fragile object as the real test

Rigid, high-friction objects are easy. Eggs, paper cups, plastic bags, wet glass and single sheets of paper are hard, because they require force control and defeat depth sensing. If a hand demo never touches one of those, the demo has told you where the limits are without saying so. We go through the physics of why in robot gripper types compared.

Watch for reorientation specifically

It’s the one dimension a casual viewer never thinks to look for and the one that most distinguishes a real hand from an expensive gripper. Make it a game: watch any manipulation video with your kid and see who spots a regrasp first.

Use benchmarking as the science-fair method

A benchmark is a complete experimental design: standardized stimulus, defined protocol, recorded outcome, reported rate. A kid who runs one has done real methodology, and judges notice the difference immediately between “I built a claw” and “I built a claw and measured it against ten standard objects over fifty trials.”

Remember that the hand is harder than the leg

Legged locomotion had a decade of visible, fundable progress. Manipulation is where the field is now stuck, because contact with varied objects is harder than contact with the floor. We lay that comparison out in why robot hands are harder than legs.

What not to do

Don’t treat degrees of freedom as a score. Thirteen is more than seven, and more is not automatically better: Boston Dynamics deleted the pinky after testing showed it unnecessary, and gained reliability by using fewer, larger actuators. A spec number without a task list attached is marketing, not measurement.

What to Watch For Over the Next 3 Months

  • Week 4: Watch for the first commercial robot hand announcement that cites a standardized benchmark result. It would be a meaningful shift in how this industry communicates, and so far nobody has done it prominently.
  • Month 2 red flags: Dexterity claims measured in degrees of freedom or finger count rather than task success. Also any hand demo where every object is rigid and textured, which is the tell that force control hasn’t been solved.
  • Month 3 self-check: Can your kid design a fair test of two grippers without being prompted, same objects, same number of attempts, success defined in advance? That’s experimental design, and it’s the most portable thing in this article.

Frequently Asked Questions

Is there one accepted dexterity score for robots?

No, and that’s the honest state of the field. There are well-established benchmark resources, the YCB object set and protocols, NIST’s task boards and competition, but no single headline number that companies report the way phone makers report battery life. Which is precisely why comparing two robot hands from their marketing is close to impossible.

Why doesn’t the Atlas hand coverage include benchmark numbers?

The October 1, 2026 IEEE Spectrum article focuses on design reasoning (degrees of freedom, actuator architecture, manufacturability) and does not specify formal dexterity benchmarks or testing protocols. That’s common in product coverage. The absence is worth noticing rather than assuming it means the testing wasn’t done.

How many degrees of freedom does a human hand have?

Roughly 20 or more depending on how you count, which is substantially more than the 13 in the new Atlas hand. But the comparison is less meaningful than it sounds, because that hand can splay its fingers beyond human range and apply forces a person can’t. Different capability envelopes, not a ranking.

What’s the single hardest manipulation task?

Thin, transparent, deformable objects: a plastic bag or a single sheet of film. Depth cameras can’t see them, suction crumples them, jaws slide off and fingers can’t locate an edge. They remain an open problem in logistics automation.

Can my kid’s project really use a research benchmark?

Yes, in spirit and partly in substance. The YCB protocols are published and readable, and household substitutes for the objects are easy to find. The valuable transfer isn’t the specific objects, it’s the discipline: define the set, define success, run a fixed number of trials, report the rate.

Does better tactile sensing fix all of this?

It would help enormously and it isn’t close. The missing piece is a sensor that is cheap, high-resolution, durable under repeated contact and easy to wire into a finger. Vision got that combination; touch hasn’t. It’s one of the clearest open hardware opportunities in the field, and a good thing to point a kid toward.


About the author

Ricky Flores is the founder of HiWave Makers and an electrical engineer with 15+ years of experience building consumer technology at Apple, Samsung, and Texas Instruments. He writes about how kids learn to build, think, and create in a tech-saturated world. Read more at hiwavemakers.com.


Sources

  1. Calli, B., Walsman, A., Singh, A., Srinivasa, S., Abbeel, P., & Dollar, A. M. (2015). “Benchmarking in Manipulation Research: The YCB Object and Model Set and Benchmarking Protocols.” IEEE Robotics & Automation Magazine, 22, 36–52. https://arxiv.org/abs/1502.03143
  2. National Institute of Standards and Technology. “Robotic Grasping and Manipulation for Assembly.” Intelligent Systems Division. https://www.nist.gov/el/intelligent-systems-division-73500/robotic-grasping-and-manipulation-assembly
  3. Ackerman, E. (2026, October 1). “Atlas Robot’s New Hand May Outperform Humanlike Designs.” IEEE Spectrum. https://spectrum.ieee.org/robust-robot-hand
  4. Ackerman, E. (2025, September 11). “Reality Is Ruining the Humanoid Robot Hype.” IEEE Spectrum. https://spectrum.ieee.org/humanoid-robot-scaling
  5. Occupational Safety and Health Administration. “Industrial Robot Systems and Industrial Robot System Safety.” OSHA Technical Manual, Section 4, Chapter 4. https://www.osha.gov/otm/section-4-safety-hazards/chapter-4
  6. International Federation of Robotics. (2026, September 24). “Five Million Robots now Operate in Factories Globally.” https://ifr.org/ifr-press-releases/news/five-million-robots-now-operate-in-factories-globally
  7. International Federation of Robotics. (2026, September 30). “Global Sales of Professional Service Robots Surge 24%.” https://ifr.org/ifr-press-releases/news/global-sales-of-professional-service-robots-surge-24-percent
Ricky Flores
Written by Ricky Flores

Founder of HiWave Makers and electrical engineer with 15+ years working on projects with Apple, Samsung, Texas Instruments, and other Fortune 500 companies. He writes about how kids learn to build, think, and create in a tech-driven world.