September 30, 2026
Humanoid

Dexterous Manipulation: Why Hands Are Still the Hardest Problem

Dexterous Manipulation Why Hands Are Still the Hardest Problem

Robotic hands lag years behind arms and legs because tactile sensing bandwidth, actuator packing density, and control complexity all hit hard physical limits at human hand scale. Despite headline demos, in-hand manipulation, repositioning an object without dropping it, remains far harder than simple grasping, and this gap will not close through better software alone.
CapabilityArms and legs (2026 state of the art)Hands (2026 state of the art)
Degrees of freedom typically achieved6-7 per arm, well-solved kinematics16-22 per hand, but many are underactuated or coupled
Sensing maturityJoint encoders and IMUs are cheap, dense, and reliableHigh-resolution tactile skin is expensive, fragile, and sparse in coverage
Task success on trained demosOften above 90 percent for locomotion and reachingFrequently 50-80 percent for contact-rich in-hand tasks
Failure recoveryFalls are rare and detectable earlySlips and drops often happen faster than sensing can react

An Opinion, Stated Plainly

The humanoid robotics industry has a messaging problem: it keeps showing off legs that walk on rubble and arms that swing boxes onto pallets, then quietly cuts away before the hand has to do anything harder than a power grip. That editing choice is not an accident. Hands are still the hardest unsolved problem in embodied robotics, and no amount of bigger vision-language-action models will paper over a fundamental mismatch between how much information a fingertip needs to sense and how little current hardware can physically deliver at that scale. This is an argument, not a neutral survey: the industry is under-investing in tactile hardware relative to how much it is investing in policy software, and that imbalance is exactly backwards for where the next real bottleneck sits.

Why Grasping Is Not Manipulation

It is worth being precise about vocabulary here, because marketing material blurs a distinction that matters enormously. Grasping is closing a hand around an object and holding it still, a power grip or pinch grip that, once achieved, requires almost no further sensing or adjustment. In-hand manipulation is different in kind: repositioning, reorienting, or regrasping an object using only the fingers, without setting it down, without using the other hand, and without dropping it. Rotating a pen between your fingers to hand someone the correct end, walking a coin across your knuckles, or adjusting your grip on a screwdriver mid-turn are all in-hand manipulation. Almost every impressive humanoid warehouse demo you have seen is grasping. Almost none is sustained in-hand manipulation, and the ones that are tend to be lab benchmarks running at a fraction of human speed and reliability.

The Sensing Bandwidth Problem

Human fingertips carry roughly 2,000 mechanoreceptors per square centimeter in the most sensitive regions, responding to vibration, shear, pressure gradients, and stretch at frequencies up into the hundreds of hertz. That is not a decorative biological detail; it is the actual information stream a human nervous system uses to detect an object beginning to slip before it has moved more than a fraction of a millimeter, and to correct grip force in single-digit milliseconds. Current tactile sensors, even the best distributed arrays being published in 2026 research, still fall well short of that combination of spatial density, temporal bandwidth, and durability at the same time. A sensor can be high-resolution or robust and cheap, but getting all three at fingertip scale, on a curved, jointed, constantly-flexing surface, remains unresolved.

The core engineering conflict is that tactile signals are spatially sparse relative to a camera image, yet temporally demanding relative to vision. A vision system can update at 30 frames per second and still catch an object moving across a table. A tactile system that only samples at 30 Hz will consistently miss the onset of a slip, because the mechanical event that precedes a dropped object unfolds in a handful of milliseconds. Covering an entire hand, not just fingertips, in sensors that are both fast enough and fine-grained enough, while surviving thousands of contact cycles against rigid and abrasive objects, is a materials science and manufacturing problem, not a machine-learning problem.

Sensing requirementHuman fingertip (approximate)Representative 2026 robotic tactile array
Spatial density~2,000 receptors per cm2 in high-sensitivity zonesTens to low hundreds of taxels per cm2 in leading research prototypes
Response bandwidthUp to several hundred Hz for vibration and slip onsetOften tens of Hz to low hundreds of Hz, with coverage tradeoffs
CoverageFull palm, fingertip, and finger-side skinFrequently fingertip-only, with palm and finger-side coverage rare
Durability under contact cyclingSelf-healing biological tissueDegradation and delamination remain open reliability problems

The Actuator Packing Problem

Sensing is only half the argument. Even with perfect tactile feedback, a hand needs somewhere to put the motors, tendons, and transmission mechanisms that generate independent finger motion, and a human-sized hand gives you almost no room to do it. A human hand has roughly 20-plus degrees of freedom driven by muscles that mostly live in the forearm, connected through tendons routed through the wrist, a packaging trick evolution arrived at because there is not enough volume in the hand itself for the actuators. Robotic hands face the identical volume constraint but without biological tendons that can be grown to fit, so engineers either route cables and motors through the wrist and forearm, mirroring the biological solution and adding mechanical complexity and friction losses, or accept fewer independently controllable degrees of freedom by coupling multiple joints to a single actuator, called underactuation, which is cheaper and more reliable but sacrifices exactly the fine independent finger control that in-hand manipulation depends on.

This is why most 16-to-22-degree-of-freedom hands you will read about are not 16-to-22 independently driven joints. Many of those degrees of freedom are coupled through linkages to a smaller number of actual motors, a compromise that keeps the hand light and reliable enough to survive real use but limits the repertoire of in-hand motions it can produce. Pushing toward fully independent actuation of every joint, which is what dexterity researchers actually want, multiplies part count, weight, heat, and failure points in a package that already has almost no spare volume.

Where the degrees of freedom actually go

A human hand has roughly 27 bones and over 20 degrees of freedom, but only a handful of intrinsic muscles live inside the hand itself; the majority of actuation is remote, routed from the forearm through tendons. Robotic hands face the same volume budget without the benefit of biological tendon material, forcing a tradeoff between independent joint actuation and mechanical simplicity.

Why Control Complexity Compounds the Problem

Even granted better sensors and better actuator packing, in-hand manipulation is a harder control problem than locomotion or reaching, for a structural reason: it is contact-rich and hybrid. Walking involves a small number of well-understood contact events, a foot touches the ground, a foot leaves the ground, and the dynamics between those events are relatively smooth and well modeled. In-hand manipulation involves multiple fingers making and breaking contact with an object and with each other, in an order that is not fixed in advance, where friction, object mass distribution, and surface compliance all change the outcome of the exact same commanded motion. Controllers have to reason about slipping contacts, rolling contacts, and sticking contacts simultaneously, often under-sensed because of the tactile coverage gaps described above.

This is also why simulation has historically underperformed for hand dexterity compared to locomotion. Physics engines model rigid-body contact reasonably well for a foot hitting a floor; they model friction-dependent, multi-point, deformable-tissue contact between a fingertip and an odd-shaped object far less reliably, which is part of why reinforcement-learning policies trained purely in simulation for dexterous manipulation historically transferred worse to real hardware than locomotion policies did. Readers interested in how the field addresses that specific transfer problem should see our companion piece on sim-to-real transfer techniques, which covers domain randomization and related methods in depth; the point relevant here is that manipulation’s contact richness makes it a harder target for those techniques than walking is.

Current Benchmarks and What They Actually Show

Recent published benchmarks such as ManiSkill-ViTac and various in-hand reorientation challenges give a more honest picture than curated demo reels. Across these benchmarks, success rates on contact-rich, occlusion-heavy manipulation tasks are meaningfully lower and far more variable than success rates on pick-and-place tasks, and performance degrades sharply when objects are novel shapes not seen during training, or when tactile sensing is removed and the policy has to rely on vision alone. Champion-level solutions in these competitions still lean heavily on task-specific engineering, careful gripper design, and constrained object sets, rather than demonstrating general-purpose dexterity across arbitrary household or industrial objects.

Common mistake

Commentary on humanoid progress regularly cites a single viral clip, a hand rotating a pen or holding an egg, as evidence that dexterous manipulation is “basically solved.” That clip almost always represents a cherry-picked run of a narrow, heavily-tuned demonstration under lab lighting and lab objects, not a benchmarked success rate across varied conditions. Judge dexterity claims by published success rates across many trials and object variations, not by the best take out of dozens of recorded attempts.

What worked

Combining sparse tactile sensing with vision-based pose estimation, rather than trying to solve manipulation from either modality alone, has produced the most robust published results. Systems that fuse a coarse but reliable tactile slip signal with a continuously updated visual estimate of object position handle occlusion and contact ambiguity better than either single-modality approach, even though neither modality alone is anywhere near human-fingertip performance.

Frequently Overlooked Details in the Dexterity Debate

  • Underactuation is a design choice, not a failureCoupling joints to fewer motors trades dexterity for reliability and weight, a reasonable choice for many industrial tasks that do not need full in-hand manipulation.
  • Tactile sensor durability is underreportedPapers showcase sensing accuracy far more often than they report how many contact cycles the sensor survives before degrading, which matters enormously for real deployment.
  • Grip force calibration driftsEven well-calibrated tactile sensors drift with temperature and wear, requiring recalibration routines that add operational overhead rarely mentioned in demos.
  • Object mass distribution is usually unknownReal objects have off-center centers of mass that a hand cannot see, only feel, making the first few hundred milliseconds of any grasp the highest-risk window.
  • Bimanual coordination hides single-hand weaknessMany demos use a second hand or a fixture to compensate for a single hand’s limited in-hand repositioning ability, masking how far single-hand dexterity actually lags.
  • Cost scales faster than degrees of freedomAdding independent actuation and tactile coverage tends to increase hand cost non-linearly, which is why most commercial humanoids ship with simpler end effectors than their research counterparts.

Why This Matters for the Humanoid Timeline

The practical consequence of this argument is that near-term commercial humanoid deployments will keep clustering around tasks that require grasping, not manipulation: moving totes, loading conveyor lines, palletizing boxes. Tasks that require sustained in-hand dexterity, fine assembly work, cable routing, tool handoffs mid-task, will remain the domain of specialized fixtures, human workers, or purpose-built end effectors for longer than optimistic roadmaps suggest. This is not a pessimistic claim about robotics broadly; arms, legs, and navigation are genuinely far along. It is a specific claim that hands are the long pole, and that treating the hand as a solved subsystem in cost and timeline projections is the single most common overreach in current humanoid forecasting.

Task categoryHand requirement2026 deployment readiness
Tote and box handlingPower grip, minimal in-hand adjustmentCommercially viable now
Bin picking of mixed rigid itemsPinch or power grip plus vision-guided regraspViable with moderate failure rates
Cable routing and small-parts assemblySustained in-hand reorientation, high tactile relianceLargely research-stage, not reliably deployed
Tool handoff and mid-task regraspDynamic in-hand repositioning under time pressureDemonstrated in constrained lab settings only

Glossary

In-hand manipulation
Repositioning, reorienting, or regrasping an object using only the fingers of one hand, without setting the object down or using the other hand.
Underactuation
A mechanical design in which multiple joints are coupled to and driven by a single actuator, reducing independent control but also reducing weight, complexity, and cost.
Taxel
A single tactile sensing element, analogous to a pixel in vision, that measures pressure, shear, or vibration at one point on a sensor array.
Slip detection
The process of sensing the earliest mechanical signs that a grasped object is beginning to move relative to the gripper, before a visible or catastrophic drop occurs.
Contact-rich manipulation
Manipulation tasks characterized by frequent, changing, and often simultaneous points of contact between fingers, the object, and the environment, as opposed to simple single-point grasping.

Key Takeaways

  • Grasping and in-hand manipulation are fundamentally different problems; most humanoid demos show the former, not the latter.
  • Human fingertips carry roughly 2,000 mechanoreceptors per square centimeter, a sensing density current robotic tactile hardware has not matched.
  • Tactile sensors face a durability, coverage, and bandwidth tradeoff that has not been solved simultaneously at hand scale.
  • Human hands route most actuation through forearm tendons because there is no room for motors inside the hand; robotic hands face the identical volume constraint.
  • Underactuated hands trade fine independent finger control for reliability and lower cost, a reasonable choice for grasping-only tasks.
  • Contact-rich manipulation is a harder simulation and control problem than locomotion, which is part of why sim-to-real transfer works less reliably for hands than for legs.
  • Near-term commercial humanoid deployments will concentrate on grasping-heavy tasks, while sustained in-hand dexterity remains a longer-term research and engineering challenge.

FAQs

What is the difference between grasping and dexterous manipulation?

Grasping is closing a hand around an object and holding it still. Dexterous, or in-hand, manipulation is repositioning, reorienting, or regrasping that object using only the fingers, without setting it down, which requires continuous sensing and control rather than a single stable grip.

Why are robotic hands behind robotic arms and legs in capability?

Hands must pack many actuators or tendons into a very small volume, rely on tactile sensors that have not matched human fingertip density and speed, and solve a harder contact-rich control problem than the relatively well-modeled dynamics of walking or reaching.

How many mechanoreceptors does a human fingertip have?

Roughly 2,000 mechanoreceptors per square centimeter in the most sensitive fingertip regions, enabling detection of slip, texture, and pressure changes at response speeds current robotic tactile sensors generally have not matched at the same spatial density.

What is underactuation in a robotic hand?

Underactuation means multiple finger joints are mechanically coupled to a single motor rather than each joint having its own independent actuator, which reduces weight, cost, and complexity but limits the range of fine, independent finger motions the hand can perform.

Why is tactile sensing harder than vision for robots?

Tactile signals are spatially sparse but must be sampled at high temporal frequency, often hundreds of hertz, to catch a slip before an object drops, while durable, high-resolution tactile arrays that survive repeated contact cycles at that speed remain difficult and costly to manufacture.

Can simulation solve dexterous manipulation the way it has helped with locomotion?

Only partially. Physics engines model rigid-body contact, such as a foot hitting the ground, reasonably well, but they model the multi-point, friction-dependent, deformable contact involved in finger manipulation less reliably, which makes sim-to-real transfer harder for hands than for legs.

What tasks are humanoid hands actually ready for today?

Power-grip and pinch-grip tasks such as moving totes, loading conveyors, and palletizing boxes are commercially viable now. Sustained in-hand tasks like cable routing, fine assembly, or mid-task tool regrasping remain largely confined to research settings.

Is dexterous manipulation mainly a software problem or a hardware problem?

It is both, but hardware constraints, tactile sensor bandwidth and durability, and the physical volume available for actuators, set a ceiling that better control software cannot fully overcome. Progress requires advances in materials and mechanical design alongside algorithmic improvements.

The policy networks that decide what a hand should do are covered in our piece on vision-language-action models. The data used to teach hands these skills in the first place is examined in teleoperation data collection. Broader workforce implications of humanoid capability gaps are discussed in our 2035 humanoid workforce analysis. For how simulation training is validated against reality more generally, see digital twins for robotics. Readers interested in whole-body coordination that complements hand dexterity can also see our piece on whole-body control.

  • arXiv, “The Role of Touch: Towards Optimal Tactile Sensing Distribution in Anthropomorphic Hands for Dexterous In-Hand Manipulation”
  • arXiv, “ManiSkill-ViTac 2025: Challenge on Manipulation Skill Learning With Vision and Tactile Sensing”
  • arXiv, “Robotic In-Hand Manipulation for Large-Range Precise Object Movement: The RGMC Champion Solution”
  • Frontiers in Robotics and AI, “TouchWGNN: Spatio-Temporal Tactile Perception for Multimodal Dexterous Manipulation”
  • arXiv, “EyeSight Hand: Design of a Fully-Actuated Dexterous Robot Hand with Integrated Vision-Based Tactile Sensors and Compliant Actuation”
  • Springer Nature, Science China Materials, “Distributed and Stretchable Tactile Sensing for Dexterous Robotic Hands Based on a Crosslinked Interpenetrating Network”
    Avatar photo
    From the University of California, Berkeley, where she graduated with honors and participated actively in the Women in Computing club, Amy Jordan earned a Bachelor of Science degree in Computer Science. Her knowledge grew even more advanced when she completed a Master's degree in Data Analytics from New York University, concentrating on predictive modeling, big data technologies, and machine learning. Amy began her varied and successful career in the technology industry as a software engineer at a rapidly expanding Silicon Valley company eight years ago. She was instrumental in creating and putting forward creative AI-driven solutions that improved business efficiency and user experience there.Following several years in software development, Amy turned her attention to tech journalism and analysis, combining her natural storytelling ability with great technical expertise. She has written for well-known technology magazines and blogs, breaking down difficult subjects including artificial intelligence, blockchain, and Web3 technologies into concise, interesting pieces fit for both tech professionals and readers overall. Her perceptive points of view have brought her invitations to panel debates and industry conferences.Amy advocates responsible innovation that gives privacy and justice top priority and is especially passionate about the ethical questions of artificial intelligence. She tracks wearable technology closely since she believes it will be essential for personal health and connectivity going forward. Apart from her personal life, Amy is committed to returning to the society by supporting diversity and inclusion in the tech sector and mentoring young women aiming at STEM professions. Amy enjoys long-distance running, reading new science fiction books, and going to neighborhood tech events to keep in touch with other aficionados when she is not writing or mentoring.

      Leave a Reply

      Your email address will not be published. Required fields are marked *