PAPER DEEP DIVE
Optimizing Grasping in Legged Robots: A Deep Learning Approach to Loco-Manipulation
This paper presents a deep learning framework designed to enhance the grasping capabilities of quadrupeds equipped with arms, with a focus on improving precision and adaptability. Our approach centers on a sim-to-real methodology that minimizes reliance on physical data collection. We developed a pipeline within the Genesis simulation environment to generate a synthetic dataset of grasp attempts on common objects. By simulating thousands of interactions from various perspectives, we created pixel-wise annotated grasp-quality maps to serve as the ground truth for our model. This dataset was used to train a custom CNN with a U-Net-like architecture that processes multi-modal input from an onboard RGB and depth cameras, including RGB images, depth maps, segmentation masks, and surface normal maps. The trained model outputs a grasp-quality heatmap to identify the optimal grasp point. We validated the complete framework on a four-legged robot. The system successfully executed a full loco-manipulation task: autonomously navigating to a target object, perceiving it with its sensors, predicting the optimal grasp pose using our model, and performing a precise grasp. This work proves that leveraging simulated training with advanced sensing offers a scalable and effective solution for object handling.
Optimizing Grasping in Legged Robots: A Deep Learning Approach to Loco-Manipulation
Paper: Optimizing Grasping in Legged Robots: A Deep Learning Approach to Loco-Manipulation
Authors: Dilermando Almeida, Guilherme Lazzarini, Juliano Negri, Thiago H. Segreto, Ricardo V. Godoy, Marcelo Becker
Affiliation: College of Mechanical Engineering, Federal University of Uberlandia, Brazil; Department of Mechanical Engineering, University of Sao Paulo, Brazil
Link: arXiv:2508.17466
One-Sentence Summary
This paper proposes a deep learning framework for quadruped robot grasping that uses the Genesis simulation environment to generate a synthetic grasp dataset, trains a U-Net architecture CNN to predict pixel-wise grasp quality heatmaps from multi-modal RGB-D inputs (RGB, depth, segmentation mask, surface normal map), and validates the complete loco-manipulation pipeline on Boston Dynamics Spot — autonomously navigating to a target object, perceiving it, predicting the optimal grasp pose, and executing a precise grasp.
Research Background and Motivation
Quadruped robots have emerged as highly efficient and versatile platforms, excelling in complex and unstructured terrains where traditional wheeled robots might fail. Equipping these robots with manipulator arms unlocks loco-manipulation capability, enabling complex physical interaction tasks in areas from industrial automation to search-and-rescue. However, achieving precise and adaptable grasping in dynamic scenarios remains a significant challenge, often hindered by the need for extensive real-world calibration and pre-programmed grasp configurations.
Traditional approaches rely on predefined grasp configurations or intensive manual calibration, limiting flexibility. Deep learning has shown promise by allowing robots to learn generalized grasping strategies from data. This paper proposes a deep learning framework specifically designed for quadruped robots, training a model to recognize objects and determine optimal grasping points from RGB-D data. Unlike most ML algorithms that return end-effector opening, orientation, and grasp quality, this work builds a pipeline integrating a pre-trained object detection model, a supervised learning model for optimal grasp point classification, and a normal estimation model.
Core contributions include: (1) a simulation-based training pipeline on Genesis Simulator enabling massive parallel data collection without real data; (2) incorporating advanced perception methods like D2NT to improve depth-based grasping precision and reduce ML execution computational cost; (3) a flexible framework integrating with high-level control APIs and commercial robots lacking low-level access; (4) validation on a physical robot demonstrating real-world effectiveness.
Synthetic Dataset Generation
Training data is entirely generated in the Genesis simulation environment from successive grasping attempts. Genesis is a simulation platform designed for physical AI/robotics/embodied intelligence, with key advantages including efficient parallel environment simulation and RGB/depth camera availability. A water bottle 3D model was selected as the grasping target.
Step 1: Virtual Camera Data Collection. The object is placed at a fixed point, and camera positions are sampled on a 2D grid of 100 X-axis points and 10 Z-axis points (both between -0.5m and 0.5m), with Y-axis fixed at $y = 0.5$ m. This yields 1000 distinct positions, each camera looking at slightly random points in X, Y (-0.03m to 0.03m), and Z (0m to 0.09m). This generates 1000 different views, each containing RGB, depth, segmentation mask, and normal map information.
Step 2: Grasp Annotation. For each pixel of the object in each image, a grasp pose is generated and labeled as successful or failed. Using the simulator segmentation mask, an end-effector is assigned to attempt grasping at each selected pixel. The pixel coordinate is converted to global frame coordinates along with the normal at that point. The end-effector is initially positioned 1.0m from the object, oriented parallel to the ground pointing toward the object, then attempts to grasp at 0.35m from the surface aligned with the normal. Success criterion: all gripper contact points collide with the bottle without any other external collision (e.g., ground contact). Grasp mask: success pixels = 1, failure = 0, outside segmentation mask = -1 (indeterminate).
Figure 1: Dataset ground truth visualization. Green pixels indicate successful grasping regions, red indicates failure, and areas outside the segmentation mask are indeterminate.
CNN Model Architecture
A fully convolutional encoder-decoder architecture (U-Net-like) performs semantic segmentation to identify optimal grasp points from multi-channel RGB-D sensor data. The encoder, based on MobileNetV2, systematically extracts features while reducing spatial resolution. The decoder uses transposed convolutions for upsampling with skip connections to merge fine-grained encoder features — critical for precise localization. The model has approximately 5.44 million trainable parameters, using GroupNorm in the decoder for training stability. A final 1x1 convolution produces a single-channel output map where the highest-value pixel indicates the best predicted grasp point.
The model is trained on a synthetic dataset generated entirely within Genesis, with virtual RGB-D camera resolution of 480x640 matching the physical sensor. The ground truth is a grasp-quality map of the same resolution, with each pixel labeled based on simulated grasp outcome: 1 for success, 0 for failure, -1 for indeterminate.
Figure 2: Model predicting optimal grasp points. Inputs (normal map, depth, segmentation, RGB) are processed by CNN to generate a grasp-quality heatmap, with the red dot indicating the optimal grasp point.
Grasping Pipeline Deployment
The complete pipeline is deployed on Boston Dynamics Spot, equipped with a 6-DOF manipulator arm and jaw gripper, with RGB camera and depth sensor on the end-effector. YOLOv11 pre-trained vision model performs object detection and segmentation, generating bounding boxes and precise binary masks.
flowchart TD A["Robot identifies object and navigates to 1m away"] --> B["Initialize gripper to open state"] B --> C["RGB-D camera captures synchronized images
480x640 resolution"] C --> D["YOLOv11 object detection and segmentation"] C --> E["D2NT normal map estimation
from depth gradients"] D --> F["Data preprocessing
normalize/align/unit normals"] E --> F F --> G["CNN predicts grasp quality heatmap"] G --> H["Select optimal pixel and compute 3D coordinates"] H --> I["Convert normal to quaternion
generate grasp pose"] I --> J["SDK sends grasp command
force-limited gripper closure"]
3D Coordinate Computation. From the selected optimal pixel $(u, v)$ and its depth value $z$, 3D coordinates are computed using camera intrinsics:
$$x = \frac{(u - u_0) \, z}{f_x}, \quad y = \frac{(v - v_0) \, z}{f_y}, \quad z = \text{depth}$$where focal lengths $f_x, f_y \approx 554.26$ pixels and principal point $(u_0=320, v_0=240)$. The normal vector is converted to Euler angles then to quaternion for manipulator control. The depth map is transformed into a normal map using the D2NT algorithm, which estimates surface normals from depth gradients.
D2NT Normal Estimation. D2NT computes surface normals directly from depth maps without constructing full 3D point clouds. For pixel $(u,v)$ with depth $z$, the 3D point is $P = (x, y, z)$. Normals are computed from local depth gradients — taking 3D point differences of adjacent pixels to build tangent planes, with the normal as the tangent plane's normal vector. The normal $\mathbf{n} = (n_x, n_y, n_z)$ is normalized to unit length:
$$\hat{\mathbf{n}} = \frac{\mathbf{n}}{\|\mathbf{n}\|} = \frac{(n_x, n_y, n_z)}{\sqrt{n_x^2 + n_y^2 + n_z^2}}$$U-Net Loss Function. The CNN is trained with pixel-level binary cross-entropy loss. For predicted grasp quality map $\hat{Y}$ and ground truth $Y$ (1 for success, 0 for failure, -1 for indeterminate):
$$\mathcal{L} = -\frac{1}{|\Omega|} \sum_{i \in \Omega} \left[ Y_i \log(\hat{Y}_i) + (1-Y_i) \log(1-\hat{Y}_i) \right]$$where $\Omega$ is the set of valid pixels within the object segmentation mask (excluding indeterminate regions), $|\Omega|$ is the count of valid pixels.
MobileNetV2 Encoder Design. The encoder uses depthwise separable convolutions from MobileNetV2, decomposing standard convolutions into depthwise and pointwise convolutions with computational cost:
$$C_{DSConv} = D_K^2 \cdot M \cdot D_F^2 + M \cdot N \cdot D_F^2$$where $D_K$ is kernel size, $M$ is input channels, $N$ is output channels, $D_F$ is feature map size. Compared to standard convolution $C_{std} = D_K^2 \cdot M \cdot N \cdot D_F^2$, this reduces computation by approximately $1/N + 1/D_K^2$, enabling efficient on-robot inference with 5.44M parameters.
Grasp Success Criterion. In the synthetic dataset, grasp success is determined by collision information. Let gripper contact point set be $\mathcal{G}$, object collision geometry be $\mathcal{O}$, and ground collision geometry be $\mathcal{E}$. The success criterion is:
$$ ext{success}(u,v) = egin{cases} 1 & ext{if } orall g \in \mathcal{G}, \, g \cap \mathcal{O} eq \emptyset \, \land \, \mathcal{G} \cap \mathcal{E} = \emptyset \ 0 & ext{otherwise} \end{cases}$$
Force-Limited Control. The SDK implements force-limited control — detecting contact force upon end-effector contact, stopping additional force upon detecting object resistance, adapting to external forces and maintaining stable grip. Maximum torque of 3.0 Nm reduces slippage or excessive force risk, accommodating minor misalignment on bottle surfaces.
Figure 3: Grasping pipeline scheme. Robot finds and walks to object → initializes RGB-D acquisition → parallel object segmentation and normal map generation → predicts grasp point → computes grasp parameters → executes grasp.
Experimental Results
The system successfully executed the complete loco-manipulation task in a real environment: autonomously navigating to the target thermos bottle, perceiving it with sensors, predicting the optimal grasp pose using the model, and performing a precise grasp. The thermos bottle has cylindrical geometry and a smooth surface, suitable for evaluating grasping performance.
| Component | Technology | Specs |
|---|---|---|
| Robot platform | Boston Dynamics Spot | Quadruped + 6DOF arm + jaw gripper |
| Object detection | YOLOv11 | Real-time detection + segmentation |
| Normal estimation | D2NT | From depth gradients |
| Grasp prediction | U-Net CNN | 5.44M params, MobileNetV2 encoder |
| Simulation | Genesis | 1000 views parallel collection |
| Camera resolution | RGB-D | 480x640 pixels |
| Max torque | Force-limited | 3.0 Nm |
| Dataset Parameter | Value | Notes |
|---|---|---|
| Camera viewpoints | 1000 | 100x10 grid sampling |
| Camera distance | 0.5m | Fixed Y-axis |
| Grasp distance | 0.35m | From object surface |
| Image resolution | 480x640 | Matches physical sensor |
| Label classes | 3 (1/0/-1) | Success/Fail/Indeterminate |
Figure 4: Processed data captured by gripper RGB-D cameras in real-world scenario. (a) RGB image (b) depth map (c) YOLOv11 segmentation mask (d) D2NT normal map (e) grasp probability heatmap (f) optimal grasping region.
Parallel simulation data generation advantage. Genesis's parallel environment capability is key to data generation. Traditional serial grasp attempts across 1000 views with hundreds of pixels each would take days. Genesis supports simultaneous parallel environments (Fig. 2 shows 5 grippers executing simultaneous grasps), drastically reducing data generation time. This scalability means adding new object categories only requires importing new 3D models and repeating the pipeline, without real robot data collection.
Sim-to-real transfer strategy. The core sim-to-real strategy: the virtual camera specs exactly match the physical RGB-D sensor on Spot's end-effector (480x640 resolution), aligning training data distribution with inference input distribution. However, real deployment still faces gaps — physical depth sensor noise and precision limitations produce lower-quality depth maps than simulation. D2NT normal estimation is sensitive to depth noise; low-quality depth maps cause normal estimation errors affecting gripper orientation. The authors partially mitigate this through preprocessing (normalizing pixel intensities to [0,1], ensuring unit normal vectors).
Limitations and Future Directions
Limitation 1: Limited model generalization. ML model generalization is limited by object geometry and texture variations, effective only for objects with specific characteristics. The training dataset contains only the water bottle model, with insufficient generalization to other shapes, sizes, and textures. The authors explicitly note the need to expand the dataset with more diverse objects and data augmentation.
Limitation 2: Depth map noise issues. Real-time RGB-D data integration faces difficulties — low quality of the end-effector depth sensor causes depth map noise, and segmentation mask resizing occasionally compromises preprocessing consistency. This reduces D2NT normal estimation and pixel-to-3D coordinate conversion accuracy. The authors suggest smoothing filters or dedicated ML models for depth map denoising.
Limitation 3: Robot SDK dependency. The pipeline currently relies on Boston Dynamics SDK tools for locomotion and manipulation control, and the lack of low-level access limits robustness in dynamic scenarios. Future work will explore locomotion and manipulation strategies beyond the SDK.
Future directions include: expanding training datasets with more object shapes/sizes/textures and data augmentation, depth map denoising (smoothing filters or dedicated ML models), exploring locomotion/manipulation strategies beyond the SDK, and testing in more complex environments with multiple objects and irregular surfaces.
Conclusion
This paper presents an innovative loco-manipulation pipeline for quadruped robots, integrating a deep learning object segmentation model with surface normal estimation. The core technical approach: leveraging Genesis simulator's parallel simulation to generate a fully synthetic grasp-annotated dataset from 1000 viewpoints (no real data needed); training a MobileNetV2-encoder U-Net CNN (5.44M parameters) to predict pixel-wise grasp quality heatmaps from RGB-D multi-modal inputs; deploying with YOLOv11 object detection, D2NT normal estimation, and CNN grasp prediction, converting optimal pixels to 3D coordinates and quaternion poses via the pinhole camera model, with force-limited control executing grasps. The complete loco-manipulation task from navigation to grasping was successfully validated on Spot. This work demonstrates the scalability and effectiveness of sim-to-real training with advanced sensing for quadruped object handling, providing a viable solution for practical applications in rescue, logistics, and domestic interaction.
SOURCE LINKS


