Skip to content
RobotWorld
Back to Papers

PAPER DEEP DIVE

灵巧手dexterous hand人形机器人

Handroid: Bridging Dexterous Hand and Humanoid

Dexterous hands and humanoid robots are typically developed as distinct embodiments: the former enable contact-rich manipulation at the object scale, whereas the latter provide mobility and whole-body interaction in human-centered environments. We introduce \textbf{Handroid}, a desktop-scale dual-embodiment robot that integrates both capabilities within a single reconfigurable platform. Handroid reuses one 27-DoF electromechanical body as either a dexterous hand or a desktop humanoid, measuring 0.33 m in height and 2.05 kg in weight. In the dexterous hand embodiment, 20 DoFs form an anthropomorphic hand closely matching the kinematic structure of the human hand. In the humanoid embodiment, the same articulated modules are reconfigured into a humanoid with a head, arms, and legs, including a 12-DoF lower-limb structure for locomotion and whole-body motion. Handroid further provides a unified control and learning framework supporting hand teleoperation, dexterous grasping, in-hand manipulation, humanoid locomotion, gait generation, and interactive motion authoring. We validate the platform through real-world dexterous manipulation, reinforcement-learning-based locomotion, keyframe motion deployment, and a long-horizon task involving embodiment reconfiguration, locomotion, docking, and dexterous pick-and-place. These results position Handroid as a compact and reproducible platform for advancing morphology-reconfigurable robotics and cross-embodiment robot learning.

Ruogu Li, Chenyang Ma, Sikai Li, Zhenyu Wei, Yunchao Yao, Haochen Shi, C. Karen Liu, Shuran Song, Mingyu DingJuly 17, 202612 min read
中文

Handroid: Bridging Dexterous Hand and Humanoid

arXiv:2607.16187 | Domain: Mechanisms & Design / Humanoid / Dexterous Hand / Robot Learning | Open-source hardware & software | 27 DoF | 0.33 m / 2.05 kg

One-Sentence Summary

Handroid uses a single 27-DoF electromechanical body that physically reconfigures between a Dexterous Hand (20-DoF anthropomorphic) and a desktop Humanoid (25-DoF with 12-DoF lower limbs) via rack-and-pinion sliding mechanisms, paired with a unified teleoperation/imitation/RL/gait-generation/motion-authoring stack, validated on real dexterous grasping (72% success), simulated gait tracking, and a cross-embodiment long-horizon task.

Background and Motivation

Morphology fundamentally defines a robot's functional boundaries: what it can reach, how it moves, what it can sense, and how it physically interacts. Dexterous hands and humanoid robots represent two complementary embodied capabilities—the former provides compact high-DoF structures for fine contact-rich manipulation (grasping, in-hand reorientation, tool use, delicate interaction), while the latter provides whole-body mobility and body-scale interaction (walking, squatting, reaching, carrying). The former emphasizes local dexterity; the latter enables global mobility.

Despite this complementarity, the two morphologies are typically developed as independent platforms. Dexterous hands are commonly mounted on fixed-base arms, providing precise manipulation within a limited workspace; humanoid robots can navigate and interact throughout an environment, but their hands are often underactuated or insufficiently dexterous for fine manipulation. Mobility and dexterity thus remain largely separated across robotic systems. This separation raises a broader question: can morphology be reused across embodiments rather than treated as a fixed property of a single robot?

Although hands and humanoid bodies differ in appearance and function, both can be viewed as collections of articulated kinematic chains connected to a central body. Fingers, arms, and legs differ in scale and task role yet share common structural elements: joint arrangements, contact geometry, actuation, sensing, and coordinated control. Handroid's core insight is that these elements can be reused across embodiments—the palm-to-finger organization is analogous to the torso-to-limb organization, so the same articulated module can serve as a finger in the hand embodiment and as an arm or leg in the humanoid embodiment.


Figure 1: Left—Handroid demonstrates dexterous manipulation in the hand embodiment and locomotion/loco-manipulation in the humanoid embodiment; Right—embodiment switching between dexterous hand and humanoid morphologies.

Hardware Design

Structure Design


Figure 2: Handroid structure design. Modules I-V correspond to the five fingers in the hand embodiment and to head, left arm, left leg, right leg, right arm in the humanoid embodiment; modules VI and VII form the base and hip. Joints 9 and 26 drive the reconfiguration sliding mechanism.

From a morphological perspective, both the human body and hand can be abstracted as branching topologies where multiple articulated chains extend from a compact central structure. The torso-to-limb organization is analogous to the palm-to-finger organization. This motivates Handroid's central design principle: the same articulated module can assume different functional roles across embodiments.

In the hand embodiment, 20 articulated DoFs are distributed across five digits, each providing one abduction-adduction DoF and three flexion-extension DoFs, approximating a commonly used 21-DoF human-hand model and supporting thumb opposition, multi-finger coordination, and in-hand manipulation. In the humanoid embodiment, 25 articulated DoFs are distributed across a 4-DoF head (Module I), two 4-DoF arms (Modules II, V), two 6-DoF legs (Modules III, IV), and a 1-DoF central hip (Module VII), with Module VI as the central torso. The remaining two actuated DoFs are prismatic joints dedicated to reconfiguration. The two legs provide 12 lower-limb DoFs approximating principal human lower-limb joint motions.

The reconfiguration mechanism is a key innovation. Two compact linear mechanisms associated with joints 9 and 26 are integrated within Module VI, each using rack-and-pinion transmission to convert actuator rotation into translation along a rigid linear guide rail. When switching from humanoid to hand, these mechanisms translate Modules II and V downward to hand-configuration positions, serving as the index- and little-finger modules—embodiment transition requires only controlled module repositioning, with no hardware replacement. An electromagnetic flange provides ~180 N holding force for rapid docking/detachment with a Franka arm.

Electrical Design


Figure 3: Electrical design—a compact mainboard integrates control, sensing, power management, and status monitoring for dual-embodiment operation.

The electrical system is designed as a compact integrated backbone, unifying actuation control, power delivery, wireless communication, onboard sensing, and hardware-state monitoring within limited internal volume. A custom mainboard coordinates all Dynamixel actuators via a TTL bus, streams robot states to a host over Wi-Fi, and supports remote operation and onboard execution of basic motion primitives. It supports both tethered development and untethered operation with onboard power and temperature monitoring. The mainboard adopts a vertically stacked modular architecture to increase usable integration area without enlarging the footprint.

Unified Control and Learning Stack

Dexterous Hand Embodiment


Figure 4: VR teleoperation as Dexterous Hand on various dexterous tasks.

Teleoperation: An Apple Vision Pro-based interface collects coordinated arm-hand demonstrations, retargeting the operator's hand motion to Handroid and mapping wrist motion to a Franka FR3 arm's end-effector command. Dexterous grasping: An object-conditioned diffusion policy is trained from demonstrations, conditioning on object geometry and recent proprioception to generate arm-hand action chunks. In-hand reorientation: An RL policy trained in simulation is deployed on the real robot, commanding hand joints from proprioceptive observations to maintain a cube in-hand while following target orientations.

Humanoid Embodiment

RL tracking control: Locomotion is formulated as reference-guided tracking. The ZMP planner alternates parameterized single/double-support phases, constructs a support-consistent desired ZMP sequence from planned footstep locations, and generates a CoM trajectory using a fixed-height linear inverted pendulum model (LIPM) with LQR preview control:

$$\ddot{x}_{\text{CoM}} = \frac{g}{z_h}(x_{\text{CoM}} - x_{\text{ZMP}})$$

where $z_h$ is the fixed height and $g$ is gravity. Mink solves inverse kinematics for leg joint angles. A closed-loop policy is trained in MuJoCo with the tracking reward:

$$r_t^{\text{track}} = 0.5\,r_t^{\text{root,pos}} + 0.5\,r_t^{\text{root,ori}} + r_t^{\text{body,pos}} + r_t^{\text{body,ori}} + r_t^{\text{body,lin}} + r_t^{\text{body,ang}}$$

RL velocity control: A reference-free policy conditioned on commanded planar CoM velocity and yaw rate outputs joint-position targets, combining velocity tracking, yaw tracking, and regularization:

$$r_t^{\text{vel}} = w_v \exp\!\left(-\frac{\|\mathbf{v}_{xy,t} - \mathbf{v}_{xy,t}^d\|^2}{\sigma_v^2}\right) + w_\omega \exp\!\left(-\frac{(\omega_{z,t} - \omega_{z,t}^d)^2}{\sigma_\omega^2}\right) + r_t^{\text{reg}}$$

Keyframe motion control: A Viser-based Keyframe Editor supports rapid whole-body motion authoring—users specify joint-space keyframes and interpolate to generate reference trajectories, enabling walking, turning, squatting, pull-ups, push-ups, and pick-and-place.

Experimental Results

Experiments address three questions: Q1 Can a reconfigurable body retain a dedicated dexterous hand's fine manipulation? Q2 Can the same body support stable locomotion and whole-body interaction as a humanoid? Q3 Can embodiment reconfiguration extend the task space beyond what either embodiment achieves alone?

Dexterous Hand Embodiment

Dexterous grasping: 10 objects of varied shapes/sizes; FoundationPose with RealSense L515 estimates 6D pose; 512 surface points sampled as policy input. 10 demonstrations per object (100 total) train the object-conditioned diffusion policy. Object poses are randomized at evaluation. Table 1 reports an average success rate of 72%.

EvaluationResult
Dexterous grasping avg success (10 objects)72%
RL tracking joint-position error0.12 rad
RL tracking body-position error0.0019 m
RL velocity tracking error (cmd 0.20 m/s)0.052 m/s

Table 1: Key quantitative results—dexterous grasping and simulated gait control both reach usable precision.

In-hand reorientation: Handroid's hand embodiment is mounted on a Franka FR3 arm, palm up, with a 3D-printed cube matching simulation dimensions. The policy runs at 30 Hz; representative rollouts are shown in Figure 6.


Figure 6: Real-world cube in-hand reorientation in the Dexterous-Hand embodiment.

Humanoid Embodiment

Simulated RL tracking: With ZMP-generated walking, the policy achieves joint-position error 0.12 rad and body-position error 0.0019 m. Simulated RL velocity control: At a commanded 0.20 m/s, the tracking error is 0.052 m/s. Keyframe motion control deploys walking, turning, squatting, pull-ups, push-ups, and pick-and-place on the real robot.


Figure 7: Humanoid embodiment tasks—pick-and-place, pull-up, push-up, etc.

Cross-Embodiment Long-Horizon Task

The final demonstration combines embodiment switching, humanoid locomotion and object interaction, electromagnetic detachment/re-docking with a Franka arm, and dexterous pick-and-place. This directly answers Q3: reconfiguration extends the task space beyond either single embodiment—the humanoid moves to the docking position, switches to the hand embodiment, and is carried by the Franka arm for dexterous manipulation.


Figure 5: Representative dexterous grasping rollouts and details.

Kinematic Consistency of Reconfiguration

A key constraint is that the same module must maintain kinematic validity in both embodiments. Let module $j$'s joint space be $\mathcal{Q}_j^{ ext{hand}}$ in the hand and $\mathcal{Q}_j^{ ext{hum}}$ in the humanoid. The rack-and-pinion translation $\Delta s$ must ensure no collision and valid joint limits in both configurations:

$$\Delta s = d_{ ext{switch}}, \quad ext{s.t.} \quad \mathcal{Q}_j^{ ext{hand}} \cap \mathcal{Q}_j^{ ext{hum}} = \emptyset ext{ (no spatial interference)}$$

This makes the transition deterministic: given an initial configuration, the post-translation state is uniquely determined, enabling automated control.

LIPM Gait Stability

The ZMP planner relies on the linear inverted pendulum model for gait stability. Under the fixed-height $z_h$ assumption, CoM dynamics linearize, and the ZMP stability condition requires the desired ZMP to remain within the support polygon $\mathcal{S}$:

$$x_{ ext{ZMP}}^d(t) \in \mathcal{S}(t), \quad orall t \in [0, T_{ ext{step}}]$$

LQR preview control minimizes a weighted cost of CoM tracking error and ZMP deviation, producing smooth stable trajectories. The 0.0019 m tracking error indicates the policy faithfully reproduces this reference.

Diffusion Policy Action Generation

The object-conditioned diffusion policy conditions on object point cloud $\mathbf{P}_{ ext{obj}}$ and recent proprioception $\mathbf{o}_{t-L:t}$ to generate future action chunks $\mathbf{a}_{t:t+H}$:

$$\mathbf{a}_{t:t+H} \sim p_ heta(\mathbf{a} \mid \mathbf{P}_{ ext{obj}}, \mathbf{o}_{t-L:t})$$

where $H$ is the action horizon and $ heta$ the denoising network parameters. Conditioning enables generalization to unseen object poses, the mechanism behind the 72% success rate.

System Architecture Diagram

flowchart TB
  subgraph HW["27-DoF Shared Electromechanical Body"]
    M["Modules I-V
hand=fingers / humanoid=head,arms,legs"] BASE["Module VI base + VII hip"] SLIDE["Rack-pinion sliding
joints 9+26 reconfiguration"] EM["Electromagnetic flange
180N dock with Franka"] end subgraph Hand["Dexterous Hand 20 DoF"] TELE["Vision Pro teleop
hand retarget+arm control"] GRASP["Object-conditioned diffusion
100 demos 72% success"] REORIENT["In-hand reorientation RL
30Hz sim-to-real"] end subgraph Humanoid["Humanoid 25 DoF"] ZMP["ZMP planner
LIPM+LQR preview"] TRACK["RL tracking policy
err 0.12rad/0.0019m"] VEL["RL velocity policy
0.20m/s err 0.052"] KEY["Keyframe editor
Viser whole-body motion"] end subgraph Long["Cross-Embodiment Long-Horizon"] SW["Embodiment switch"] LOCO["Humanoid locomotion"] DOCK["Electromagnetic dock/detach"] DEX["Dexterous pick-and-place"] end HW --> Hand HW --> Humanoid SLIDE --> SW Hand --> DEX Humanoid --> LOCO EM --> DOCK SW --> LOCO --> DOCK --> DEX
EmbodimentDoFKey Module MappingCore Capability
Dexterous Hand20 DoFModules I-V → five fingersGrasping/reorientation/fine manipulation
Desktop Humanoid25 DoFI→head, II/V→arms, III/IV→legs, VII→hipWalking/turning/squatting/pull-up/push-up
Reconfig Mechanism2 DoF (prismatic)Joints 9+26 → rack in Module VISwitching without hardware replacement

Table 2: DoF allocation and module mapping across dual embodiments—27 DoF redistributed; reconfiguration occupies 2 DoF independently.

From a design philosophy perspective, Handroid's modularity echoes the biological concept of "homology." The human hand and foot are highly homologous in skeletal structure—both composed of phalanges, metacarpals/tarsals, carpals—yet differentiate into fine manipulation and weight-bearing locomotion. Handroid's modules I-V serving as fingers in one embodiment and head/arms/legs in another is an engineering realization of this homology. This design not only saves hardware cost but lets the same actuation, sensing, and control interfaces be reused across embodiments, providing a unified abstraction for cross-embodiment learning.

The cross-embodiment long-horizon task deserves deeper interpretation. The robot first moves in humanoid form to a docking position, docks with a Franka arm via the electromagnetic flange, switches to hand form, and is then carried by the arm for dexterous pick-and-place. This demonstrates a "mobile-dock-manipulate" cooperative paradigm: the humanoid provides mobility to reach the workspace, the hand provides fine manipulation, and the external arm provides large-workspace positioning. This paradigm maps real-world scenarios like "inspection robot reaches equipment, docks an arm, performs fine maintenance"—demonstrated at desktop scale but generalizable.

Limitations

Early-stage platform (author-acknowledged). Handroid is an initial step toward dual-embodiment robot learning. Future wireless operation will reduce cable-induced disturbances and enable longer mobile tasks. Further miniaturization of actuators and structures can improve the hand embodiment while preserving humanoid functionality.

Lack of tactile sensing (author-acknowledged). Fingertip cameras and tactile sensors are not yet integrated. Adding them would benefit both contact-rich manipulation and foot-ground contact estimation—since the same modules serve as both fingers and feet.

Desktop-scale limits task scope. The 0.33 m / 2.05 kg desktop scale aids low-cost reproducibility and rapid iteration but limits direct interaction with true human-scale environments. Whether morphology reuse maintains mechanism efficiency and kinematic plausibility at full-body scale remains to be validated on larger platforms.

Cross-embodiment policy transfer not yet explored. The paper demonstrates per-embodiment policies, but cross-embodiment transfer (e.g., grasping knowledge from the hand embodiment informing arm manipulation in the humanoid embodiment) is mainly proposed as future work, not validated experimentally. This is the deepest value of the morphology-reuse idea and the most open challenge.

Summary and Outlook

Handroid realizes physical reconfiguration between dexterous hand and humanoid embodiments with a single 27-DoF body and a unified control/learning stack, validating morphology reuse on real grasping (72%), simulated gait tracking (0.12 rad / 0.0019 m), and a cross-embodiment long-horizon task. The rack-and-pinion sliding mechanism enables embodiment switching without hardware replacement; the electromagnetic flange enables rapid docking/detachment with an external arm.

From a broader perspective, Handroid represents a paradigm shift from "one robot, one morphology" to "one robot, multiple morphologies." When the same module can be both a finger and a leg, morphology is no longer a fixed property but a configurable degree of freedom. This not only lowers hardware cost but provides a unique experimental platform for cross-embodiment learning—the same physical structure learns different skills in different embodiments, and transfer/sharing between them may reveal deeper principles of robot learning. With tactile sensing and wireless operation, Handroid could become a standard research platform for morphology-reconfigurable robotics.

Golden Lines

"Can morphology be reused across embodiments rather than treated as a fixed property? The palm-to-finger organization is analogous to the torso-to-limb organization."
"When the same module serves as both finger and foot, morphology is no longer a fixed property but a configurable degree of freedom."

Related Papers

Pre-training Visual Dexterity in Simulation

Pre-training Visual Dexterity in Simulation

Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers. Dexterous, multi-fingered hands remain comparatively data-starved because real teleoperation is costly to scale, while human hand video is off-embodiment and requires lossy pose estimation and retargeting. We introduce Simulation Pre-training for Dexterity (SPD), a pre-training framework for dexterous manipulation that uses data entirely collected in simulation. In SPD, humans manipulate virtual objects inside a VR headset, enabling on-embodiment trajectories and robot-free collection. With the help of five operators, we collect 75 hours of multi-task dexterous manipulation over one week, and use it to pre-train a causal transformer on a sequence modeling objective. We study the benefits of simulation pre-training on real-world tasks by fine-tuning on 1-2 hours of physical demonstrations on a 56-DoF bimanual dexterous setup. We find that our approach outperforms training behavior cloning policies from scratch, showing that simulation teleoperation is a viable pre-training source for real-world dexterous manipulation. We perform ablation studies, measuring the benefits of history conditioning and short action chunks for reactive control.

灵巧操作灵巧手预训练Aug 16, 2026
VTAP Gripper: Synergizing Fingertip Sensing and a Visuo-Tactile Active Palm for Dexterous In-Hand Manipulation

VTAP Gripper: Synergizing Fingertip Sensing and a Visuo-Tactile Active Palm for Dexterous In-Hand Manipulation

This paper presents a tactile-reactive gripper that integrates a Visuo-Tactile Active Palm (VTAP) and compliant, reconfigurable fingers equipped with tactile array sensors. The design exploits structured finger-palm synergy and multi-modal perception to achieve both robust grasping and fine manipulation. The actuated bi-modal palm seamlessly combines long-range visual localization with contact-rich tactile feedback, substantially extending the system's manipulation capability. To bridge the embodiment gap between human hand motion and the heterogeneous three-finger structure, we further propose a staged, gesture-conditioned retargeting framework for dexterous teleoperation. Extensive experiments validate the system across a range of challenging tasks: reactive grasping of YCB and fragile objects, in-hand syringe reorientation and plunger actuation, singulation of clustered objects down to 3 mm in diameter, and vision-tactile peg-in-hole insertion. Results demonstrate that high manipulation performance can be achieved through coordinated finger-palm interaction and multi-modal sensing, without resorting to high degrees of freedom anthropomorphic designs. The VTAP gripper and its retargeting framework offer a practical reference architecture for dexterous gripper design, manipulation, and contact-rich data collection in support of learning-based approaches. Project webpage: https://yuhochau.github.io/vtap/.

灵巧手触觉传感手内操作Jul 16, 2026
PAPERdexpoint-2024DexPoint:Generalizable PointCloud Policies for灵巧手点云操作

DexPoint: Generalizable Point Cloud Policies for Dexterous Manipulation

DexPoint uses point cloud representations for generalizable dexterous manipulation without CAD models, showing strong sim-to-real transfer.

灵巧手点云操作