Function beyond Form: Functional Correspondence for Cross-Embodiment Dexterous Grasp Generation

Bolin Zou1, Wenlong Dong1, Mu Ai1, Chao Tang2, Aoxiang Gu1, Lipeng Chen3,4, Hong Zhang1

1 Shenzhen Key Laboratory of Robotics and Computer Vision, Southern University of Science and Technology, Shenzhen, China.
2 Department of Robotics, Perception and Learning, KTH Royal Institute of Technology, Stockholm, Sweden.
3 School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China.
4 Rysen Robotics, Shenzhen, China.

Project videoLoading video player…

Abstract

Cross-embodiment dexterous grasp generation remains challenging because robotic hands differ substantially in geometry, topology, and kinematics. Existing approaches often lack explicit correspondences between structurally different hand regions that play similar functional roles in a grasp, a concept we refer to as functional correspondence. Consequently, their models tend to learn hand-specific interaction patterns rather than transferable grasp knowledge, limiting generalization to unseen hands. To address this limitation, we introduce FunCo-Grasp, which establishes functional correspondences across heterogeneous hand embodiments. Specifically, Functional Part Alignment aligns each hand to a canonical functional schema by mapping physical links to shared functional parts according to their grasping roles, while Canonical Frame Alignment expresses these parts in canonical local frames. These two alignments provide a consistent representation for inter-part and hand-object interactions, allowing the model to learn transferable grasp knowledge across hands. Conditioned on the aligned hand representation and object geometry, a diffusion model generates the target spatial arrangement of the functional parts, which are then converted into an executable joint configuration. Adapting FunCo-Grasp to an unseen hand requires only its geometric and kinematic models and a one-time lightweight functional annotation, without target-hand grasp data, fine-tuning, or learned retargeting. In simulation on held-out objects from the filtered CMapDataset, FunCo-Grasp achieves average success rates of 92.40% on three seen hands and 74.02% on four unseen hands. In real-world experiments, the same model achieves an overall success rate of 76.00% on two unseen hands without additional training or fine-tuning. These results demonstrate the effectiveness of FunCo-Grasp in transferring grasp knowledge to unseen hands.

Pipeline Overview

Functional parts, canonical local frames, and cross-embodiment grasp generation.
Overview of FunCo-Grasp. Top: The canonical functional schema and cross-embodiment functional correspondences established through a one-time lightweight functional annotation via (a) Functional Part Alignment and (b) Canonical Frame Alignment. Bottom: Grasp transfer from seen to unseen hands.
FunCo-Grasp pipeline overview.
Overview of FunCo-Grasp. (1) Cross-Embodiment Functional Correspondence maps physical links to shared functional parts according to their grasping roles and expresses them in canonical local frames. (2) Hand and Object Encoding extracts node features N, edge features E, and object patch features O. (3) Functional-Part-Based Grasp Generation conditions on the aligned hand representation and object geometry to generate the target spatial arrangement of the functional parts in the shared interaction space, before converting them into executable joint configurations via inverse kinematics.

Simulation Results

Generated grasps on held-out objects across seen and unseen embodiments.
MethodSuccess rate (%) ↑Eff. (s) ↓
Seen embodimentsUnseen embodiments
AllegroBarrettShadowHandAvg.ApexHandXHandLEAPRobotiq-3FAvg.
GenDexGrasp67.0050.0053.3056.7715.1729.3725.3331.0025.2218.77
CEDex92.6685.1686.4388.0846.2062.0071.5086.3066.5012.52
D(R,O) Grasp89.9088.3077.8085.3324.2058.604.8064.9038.130.75
T(R,O) Grasp89.1091.6092.5091.071.304.000.1042.4011.950.21
UniMorphGrasp†90.3093.0098.8094.00----------0.45
FunCo-Grasp (Ours)97.2086.2093.8092.4068.2072.7380.4074.7374.020.16

Grasp success on held-out objects and per-grasp runtime. Red and yellow mark the best and second-best values. †UniMorphGrasp results come from its paper because its code is not publicly available. -- denotes unavailable results.

Real-World Experiments

ApexHand

ApexHand real-world experimentsLoading video player…

Revo2

Revo2 real-world experimentsLoading video player…