PAPER DEEP DIVE
Learning Dynamic Pick-and-Place for a Legged Manipulator
Legged manipulators extend robotic capabilities beyond static manipulation by integrating agile locomotion with versatile arm control. However, achieving precise manipulation while maintaining coordinated locomotion remains a major challenge. This work presents a hierarchical reinforcement learning framework for dynamic pick-and-place tasks using a quadruped equipped with a 6-DOF robotic arm. The framework incorporates an explicit mass estimation module enabling adaptive whole-body control for objects with varying weights. In simulation, the system achieves an 86.05% success rate with payloads up to 2.3 kg. The approach is further validated through real-world experiments across six representative scenarios with controlled variations in object physical properties (size and mass) and task heights. Specifically, within a wide vertical workspace ranging from ground level to 1.1~m-high tabletops, the system demonstrates an average success rate of 73.3% for payloads up to 1.3 kg, with an average execution time of 4.06 s. Unlike prior works that handle lightweight objects and execute pick-and-place motions with slow, piecewise motions, the proposed framework exploits concurrent locomotion and manipulation for dynamic, continuous execution. These results demonstrate the potential of quadrupedal mobile manipulators for adaptive, whole-body pick-and-place with heavier payloads and extended workspaces.
Learning Dynamic Pick-and-Place for a Legged Manipulator
Paper: Learning Dynamic Pick-and-Place for a Legged Manipulator
Authors: Moonkyu Jung, Jiseong Lee, Zhengmao He, Donghoon Youm, Juhyeuk Mun, HyeongJun Kim, Hyunsik Oh, Donghyuk Choi, Jungwoo Hur, Jie Song, Jemin Hwangbo
Institutions: KAIST Robotics and AI Lab, HKUST(GZ) Robotics and Autonomous Systems
Links: arXiv:2605.15713
One-Sentence Summary
This paper presents a hierarchical RL framework enabling a quadruped + 6-DoF arm legged manipulator to perform dynamic continuous pick-and-place — a low-level controller stabilizes whole-body motion under random arm disturbances, while a high-level controller uses LSTM-based online mass estimation for adaptive whole-body control, achieving 86.05% simulation success rate (payload up to 2.3 kg) and 73.3% real-world average across 6 scenarios (payload up to 1.3 kg) in just 4.06s average — 10× faster than prior quasi-static approaches.
Background and Motivation
Legged manipulators — robots integrating a quadrupedal base with a robotic arm — have the potential to extend pick-and-place capabilities beyond the structured, planar environments constraining traditional fixed-base and wheeled manipulators. Quadrupedal bases provide high DoF and adaptive mobility, enabling diverse dynamic behaviors from stable walking to agile parkour. This work designs a controller that fully exploits the quadruped's locomotion and whole-body posture for fast, accurate pick-and-place.
Existing methods are limited by quasi-static execution — stopping before manipulating, decomposing pick-and-place into slow, piecewise motions. Yokoyama et al. decomposed the task into modular visual-motor skills but relied on the robot's internal controller, resulting in slow execution without unified locomotion-manipulation coordination. Zhang et al. distilled multi-stage teacher policies into a single student policy, achieving vision-based autonomous pick-and-place but with quasi-static execution averaging 43.8s per task, largely limited to floor-level placement. These methods involve discrete "stop-then-manipulate" transitions, unable to exploit concurrent locomotion and manipulation.
The key insight: through hierarchical policy structure (low-level base stabilization + high-level task execution) with an explicit mass estimation module for adaptive whole-body control, the legged manipulator can perform precise grasping while the base remains in motion — the end-effector velocity approaches zero at contact while the base maintains up to 1 m/s planar speed and -1 rad/s yaw rate. This dynamic coordination reduces execution time from 43.8s to 4.06s (10× speedup) and extends the workspace from ground level to 1.1 m-high tabletops — more than twice the robot's base height.
Method
The system uses a hierarchical control structure. The low-level controller generates desired positions for 12 leg joints to stabilize base motion and regulate whole-body posture; the high-level controller generates task-level commands controlling both locomotion and manipulation. The high-level action decomposes into base and arm components:
$$\boldsymbol{a}^{\text{high}}=[a^{\text{base}},a^{\text{arm}}]$$
where base command $a^{\text{base}}=[v_{x},v_{y},\omega_{z},\phi,\Delta h]$ includes linear/angular velocities, pitch, and height change, and arm command $a^{\text{arm}}=[q^{\text{arm},0}_{\text{des}},\ldots,q^{\text{arm},5}_{\text{des}},a^{\text{gripper}}]$ includes 6 joint positions and gripper open/close.
Figure 1: Hierarchical training framework. Step 1: a low-level locomotion controller is trained via RL for robust whole-body motion; Step 2: a high-level controller is trained for pick-and-place on top of the low-level controller.
Low-level Whole-body Locomotion Controller. Receives base velocity command $[v_{x},v_{y},\omega_{z}]$ and body-control command $[\Delta h,\phi]$, outputs desired positions for 12 leg joints $q^{leg}_{des,t}$. Policy observation $o^{\text{low}}=[x,d,a]$ includes base/leg proprioceptive state $x$, arm proprioceptive state $d$, and previous action $a$. Uses asymmetric actor-critic with privileged critic observations $[h_{base},V_{base},T_{air},T_{stance},f_{grf},H_{scan}]$.
Unlike prior methods that uniquely determine base posture from the end-effector target in a ground-fixed frame, this work treats base posture as a controllable variable, allowing the high-level controller to flexibly search for optimal base and arm configurations.
Random Arm Motion Generator. Applies dynamic disturbances during low-level training. Each joint independently selects target position and trajectory duration, producing asynchronous motion patterns. Joint displacement:
$$\Delta q_{i}=q_{\text{target},i}-q_{\text{init},i}$$
Each joint $i\in\{0,1,2,3,4,5\}$ independently chooses constant-velocity or symmetric constant-acceleration motion. A random payload up to 2 kg is attached during training. The controller implicitly learns to compensate for external forces from arm state histories.
High-level Controller. Observation $o^{\text{high}}=[s^{r},s^{o},a^{\text{high}},p_{\text{pick}},p_{\text{place}},i_{\text{phase}}]$ includes robot state, object SE(3) state, previous action, pick/place positions, and phase index. The key component is an LSTM-based online mass estimation module that estimates object mass $m_{\text{obj}}$ and contact state $C_{\text{contact}}$ from history, enabling adaptive lift height adjustment — proactively increasing lift margin for heavier objects.
flowchart TD A[High-level obs o_high] --> B[LSTM Mass Estimation] B --> C[Object mass m_obj] B --> D[Contact state C_contact] C --> E[Adaptive whole-body control] D --> E E --> F[Base command a_base] E --> G[Arm command a_arm] F --> H[Low-level controller] G --> H H --> I[12 leg joint positions] H --> J[6 arm joints + gripper] I --> K[Quadruped base] J --> L[Robot arm] K --> M[Dynamic grasp: base motion + precise arm control] L --> M M --> N[Continuous pick-and-place 4.06s]
Success-rate-driven Curriculum. Three curriculum levels $L_{\text{pick}}$, $L_{\text{place}}$, $L_{\text{release}}$ initialize near zero and increment toward 1.0 based on subgoal success rates. Difficulty is controlled through (i) robot-to-pick-table distance, (ii) inter-table distance, (iii) object mass, (iv) place/release success criteria. Every 50 iterations, each level increases by 0.01 if success rate exceeds threshold (pick 0.90, place 0.85, release 0.85). A synchronization constraint keeps $L_{\text{pick}}$ and $L_{\text{place}}$ within 0.015 margin.
Reward Design. The task decomposes into 6 stages — pre-grasping, grasping, carrying, placement, retreating, finishing — each combining sparse and dense rewards. The pre-grasping dense reward guides the end-effector toward the object:
$$R_{\text{EE\_to\_Obj}}=k_{1}\cdot e^{-25\cdot\|p_{\text{obj}}-p_{\text{ee}}\|^{2}}+k_{2}\cdot e^{-1\cdot\|p_{\text{obj}}-p_{\text{ee}}\|^{2}}$$
The dual-exponential term balances precise approach (σ=1/25) with long-range guidance (σ=1). The carrying-stage object-to-place dense reward uses a similar dual-exponential:
$$R_{\text{Obj\_to\_Place}}=k_{6}\cdot e^{-5\cdot\|p^{key}_{\text{obj}}-p^{key}_{\text{place}}\|^{2}}+k_{7}\cdot e^{-25\cdot\|p^{key}_{\text{obj}}-p^{key}_{\text{place}}\|^{2}}$$
Placement requires the object to be upright, penalizing z-axis deviation:
$$R_{\text{orient}}=k_{27}\cdot e^{-\|\mathbf{p}_{\text{obj}}-\mathbf{p}_{\text{place}}\|}\cdot(\mathbf{R}_{\text{obj}}^{(z)}-1)^{2}$$
This penalty increases as the object approaches the placement point, ensuring upright placement. Base penalty regularizes vertical velocity and angular velocities:
$$R_{\text{base\_pen}}=k_{19}\cdot(V^{2}_{z}+0.02\cdot|\omega_{x}|+0.02\cdot|\omega_{y}|)$$
Results
Experiments use an in-house quadruped + Unitree Z1 arm (6-DoF arm + 1-DoF gripper, ~2.3 kg effective payload). Success requires placing the object upright within 5 cm of place table center and retreating ≥10 cm within 10s.
| Method | Success (%) | Time (s) | Notes |
|---|---|---|---|
| Ours (explicit est.) | 86.05 | 3.92±0.40 | Unified policy + LSTM mass est. |
| Latent adaptation | 85.5 | 3.98 | Dual encoders, higher cost |
| w/o estimation | 82.0 | 4.05 | Domain randomization only |
| Segmented policy | 70.0 | 5.26 | No value propagation across phases |
The explicit estimation module matches latent adaptation performance with 11.32% less training time. The w/o-estimation baseline degrades more for out-of-distribution heavy objects — explicit mass estimation enables proactive lift height adjustment. The segmented policy suffers significant success rate drop due to myopic grasping that ignores placement suitability.
Figure 2: Low-level controller disturbance rejection. Yaw-velocity fluctuation under periodic Joint 1 motion (1 rad amplitude) is effectively suppressed.
| Scenario | Sim Success | Real Success | Real Time (s) |
|---|---|---|---|
| Nominal (∅6×10cm, 0.83kg, 0.6↔0.9m) | 9/10 | 9/10 | 4.07±0.66 |
| Heavy (1.3 kg) | 8/10 | 3/10 | 4.76±0.86 |
| Light (0.47 kg) | 8/10 | 8/10 | 3.74±0.72 |
| Square (5×5×10cm) | 9/10 | 8/10 | 4.08±0.74 |
| Large (∅7×10cm) | 10/10 | 6/10 | 3.74±0.97 |
| Large height gap (0↔1.1m) | 7/10 | 7/10 | 3.54±0.61 |
Real-world average across 6 scenarios: 73.3% success, 4.06s average. The heavy object scenario degrades most (sim 8/10→real 3/10), likely due to Z1 arm torque tracking degradation under 1.3 kg load. Large-size objects also degrade (10/10→6/10) as larger diameter makes grasping more prone to slipping.
Figure 3: Base motion analysis at grasping moment. End-effector velocity approaches ~0.2 m/s while base maintains 1 m/s planar speed and -1 rad/s yaw rate, achieving dynamic coordination.
Dynamic Grasping Analysis
The motion analysis at the grasping moment is the core highlight. In the fastest cases, the end-effector velocity smoothly converges to ~0.2 m/s while the base maintains up to 1 m/s planar speed and -1 rad/s yaw rate. This demonstrates true dynamic coordination: the base provides continuous momentum for efficient task execution while the arm executes the precise, low-velocity control required for reliable grasping — fundamentally different from quasi-static methods that must fully stop the base before grasping.
Limitations
Heavy object real-world degradation. The 1.3 kg scenario achieves only 30% real-world success (vs. 80% sim), as Z1 arm torque tracking degrades under heavy load. Training uses up to 2.3 kg but real-world validation only reaches 1.3 kg.
Perception simplification. This work focuses on dynamic capability rather than perceptual sophistication; object poses are assumed known in real-world experiments. Real deployment requires integrated visual perception (depth camera + object detection), adding latency and error.
Conclusion
This paper presents a hierarchical RL framework for dynamic pick-and-place with a legged manipulator. The low-level controller stabilizes the base under random arm disturbances; the high-level controller uses LSTM mass estimation for adaptive whole-body control. Six-stage reward design and success-rate-driven curriculum enable stable convergence of the long-horizon task. Simulation achieves 86.05% success; real-world 6-scenario average is 73.3% in 4.06s — 10× faster than quasi-static methods, with workspace from ground to 1.1 m-high tabletops. Dynamic grasping maintains 1 m/s base motion while the end-effector approaches zero velocity, achieving true concurrent locomotion and manipulation.
SOURCE LINKS

