PAPER DEEP DIVE
Robot Learning to Communicate through Projected Visual Abstractions
Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.
Robot Learning to Communicate through Projected Visual Abstractions
Paper: Robot Learning to Communicate through Projected Visual Abstractions
Authors: Danyang Yan, Boyuan Wang, Jiaxun Liu, Boyuan Chen
Affiliations: Duke University (ECE / MEMS / CS)
Links: arXiv:2607.22434 · Project Page
One-Sentence Summary
This paper presents a robotic system for dynamic shadow expression using a 21-DoF soft-skinned dexterous hand and a differentiable shadow self-model that learns the mapping from hand configurations to projected shadow appearance through task-agnostic self-exploration, optimizing hand configurations via gradient search and refining through collision-aware simulation to obtain physically feasible motions, validated across sign-language gestures, hand-shadow puppetry, and animal motion imitation.
Background and Motivation
Humans routinely communicate through abstractions of their bodies — shadows, silhouettes, and reflections. These representations allow observers to infer information about objects, people, actions, and even emotions without directly observing the physical entity. Yet robots remain largely confined to expressing themselves through physical morphology. Enabling robots to communicate through projected visual abstractions requires reasoning about both bodily motion and how that motion transforms into an external representation perceived by an observer.
Projected visual abstractions are visual forms that emerge from but are distinct from the physical embodiment. Shadow performance is particularly intriguing because shadows arise directly from the geometry and motion of a body while transforming it into a fundamentally different visual representation. The core challenge: robots must generate shadows that are both visually recognizable and physically controllable. Conventional rigid hands are poorly suited — gaps between rigid linkages allow light leakage, fragmenting silhouettes. Fully soft hands lack precise motion control. Human hand-shadow artists achieve this through rigid skeletal support covered by soft tissue.
Figure 1: A robot hand actively generating expressive shadow performances.
Method
1. Hybrid Rigid-Soft Hand Design
Figure 2: Robot hand design. Rigid skeleton covered by compliant soft skin; each finger is a four-link kinematic chain with 2-DoF MCP joints.
Each finger is a slender rigid skeleton covered by a compliant foam layer fabricated via low-infill TPU 3D printing. Each finger is a four-link kinematic chain approximating DIP, PIP, and MCP joints. Each MCP joint provides 2 DoF (flexion-extension and abduction-adduction), plus a wrist joint, totaling 21 DoF. Projection setup: spotlight 3m from palm center, white backdrop 1m behind the hand.
2. Shadow Self-Model Learning
Figure 4: Shadow self-modeling framework. (A) Robot collects joint-shadow pairs via motor bubbling; (B) At deployment, model frozen, gradients optimize joint angles to match target.
The mapping from joint configurations to shadow appearance is difficult to invert — many physically distinct hand poses produce identical or highly similar 2D silhouettes, making direct inverse model training ill-posed. A forward shadow self-model is trained to predict binary shadow images $\hat{\mathbf{I}}_i$ from 21-D joint configurations $\bm{\uptheta}_i$:
$$ \hat{\mathbf{I}}_i = f_{\text{NN}}(\text{FK}(\bm{\uptheta}_i)), \quad \mathcal{L}_{\text{MAE}} = \|\hat{\mathbf{I}}_i - \mathbf{I}_i\|_1 $$The forward kinematics maps joint angles to fingertip positions; for finger $j$:
$$ \mathbf{p}_j = \text{FK}(\bm{\uptheta}_j) = \mathbf{T}_{\text{wrist}} \prod_{k=1}^{4} \text{Exp}(\bm{\uptheta}_{j,k}) \, \mathbf{p}_{\text{tip}} $$where $\bm{\uptheta}_{j,k}$ is the $k$-th joint angle of finger $j$ and $\text{Exp}(\cdot)$ is the exponential coordinate transform. Data is collected self-supervised via motor bubbling.
3. Two-Stage Optimization
Stage 1: gradient optimization on the frozen self-model. Total loss combines pixel reconstruction, silhouette overlap, and perceptual similarity:
$$ \mathcal{L}_{\text{total}} = w_1 \cdot \mathcal{L}_{\text{MAE}} + w_2 \cdot \mathcal{L}_{\text{IoU}} + w_3 \cdot \mathcal{L}_{\text{CLIP}} $$Stage 2: hill-climbing refinement in a physics simulator with collision checking. 39.34% of self-model-optimized poses have collision issues. The hill-climbing search finds collision-free configurations near the self-model solution:
$$ \bm{\uptheta}^{\text{final}} = \arg\min_{\bm{\uptheta}} \mathcal{L}_{\text{total}}(\bm{\uptheta}) \quad \text{s.t.} \quad \text{collision-free}(\bm{\uptheta}), \quad \|\bm{\uptheta} - \bm{\uptheta}^*\| \leq \epsilon $$4. Dynamic Shadow and Motion Planning
For dynamic shadow performance, keyframes are extracted from target videos and optimized independently with temporal smoothness regularization:
$$ \mathcal{L}_{\text{temporal}} = \sum_{t} \|\bm{\uptheta}_{t+1} - \bm{\uptheta}_t\|_2^2 + \lambda_e \cdot \mathcal{L}_{\text{expressive}}(\bm{\uptheta}_t) $$A sample-based trajectory planner generates collision-free motion plans executed on the physical robot.
flowchart TD
A["Target shadow image/video"] --> B["Binarization"]
B --> C["Stage 1: Gradient optimization
on frozen self-model"]
C --> D["Best joint config
(may collide)"]
D --> E["Stage 2: Hill-climbing refinement
collision-aware simulation"]
E --> F["Physically feasible pose"]
F --> G["Motion planning
collision-free trajectory"]
G --> H["Physical hand execution
real shadow projection"]
I["Motor bubbling self-exploration"] --> J["Forward shadow self-model
FK + NN"]
J --> C
style C fill:#e1f5fe
style E fill:#fff3e0
style J fill:#e8f5e9
Experiments
Evaluation Setup
The dataset contains 61 targets: 26 ASL hand gesture images, 19 hand-shadow animal keyframes (duck, deer, peacock, camel), 16 raw animal video keyframes (howling wolf, raven flapping). Difficulty increases: ASL most aligned with hand morphology, hand-shadow animals more stylized, raw animals don't correspond to any hand embodiment.
Figure 5: Shadow imitation results for static image and video frame targets.
| Method | Total↓ | Base↓ | CLIP↓ | MAE↓ | IoU Loss↓ |
|---|---|---|---|---|---|
| Random | 0.4590 | 0.2135 | 0.1116 | 0.3402 | 0.4772 |
| Inverse | 0.4262 | 0.1959 | 0.1246 | 0.3102 | 0.4370 |
| Nearest Neighbor | 0.2014 | 0.0849 | 0.0782 | 0.1160 | 0.1910 |
| Ours (w/o H.C.) | 0.1809 | 0.0828 | 0.0681 | 0.1090 | 0.1873 |
| Ours (Base) | 0.1210 | 0.0640 | 0.0707 | 0.0829 | 0.1440 |
| Finding | Details |
|---|---|
| Forward > inverse model | Inverse model near random; forward self-model provides differentiable mapping |
| H.C. refinement +33.1% | Total loss 0.1809 → 0.1210, resolves 39.34% collision poses |
| Hybrid > pure hill-climbing | 500-step hybrid > 500-step pure H.C., approaches 2000-step pure H.C. with less runtime |
| Cross-embodiment generalization | ASL easiest, hand-shadow animals moderate, raw animals hardest but still meaningful |
Figure 6: Quantitative results table for all 61 single-image targets.
Figure 3: Shadow self-modeling framework details (continued).
Limitations
- Hand morphology constraints: The 21-DoF hand, though more dexterous than most, still differs from human hands; some complex shadows requiring extremely fine fingertip control may not be accurately reproduced.
- Single light source assumption: The system assumes a single spotlight 3m away with fixed projection setup; robustness to multi-source lighting, ambient interference, or varying projection distances is unverified.
- Insufficient real-time performance: Two-stage optimization (gradient search + hill-climbing) is computationally expensive, unsuitable for real-time interactive shadow dialogue scenarios.
Conclusion and Outlook
This paper establishes a framework for robots to communicate through projected visual abstractions. The hybrid rigid-soft hand design balances kinematic precision and light sealing; the differentiable shadow self-model learns the configuration-to-shadow mapping self-supervised; two-stage optimization balances differentiable efficiency with physical feasibility. Validation across sign language, hand-shadow puppetry, and animal motion imitation expands robotic expression into the domain of projected visual abstractions.
Golden quote: "Directly training an inverse model from target shadow to joint angles is ill-posed — because many physically distinct hand poses produce identical 2D silhouettes; training a forward self-model to make it differentiable, then optimizing joint angles via gradient backpropagation, is the key to translating projected appearance into embodied execution."
SOURCE LINKS