PAPER DEEP DIVE
Temporal Visuo-Tactile Learning for Dexterous Grasp Stability
Most robotic grasping research picks a pose from vision and closes a parallel gripper, while humans rely on fingertip touch and almost never drop things. The LASR Lab at TU Dresden reframes the question as "is this multi-fingered grasp stable before we lift?" and answers it with data. They mount four Digit 360 sensors on the fingertips of a 16-DoF Tilburg Hand on a 7-DoF xArm7 and automatically collect 10,000 grasp trials over 200 objects (105 rigid, 95 deformable, 1.9-244 g), recording external RGB-D, proprioception and four tactile streams (camera, audio, IMU, pressure) throughout each grasp. Stability labels come from SAM 3 measuring post-lift object height and agree with manual verification on 92.34% of trials. The predictor consumes a 3-second window anchored at grasp initiation: 16-frame RGB and tactile clips go through pretrained VideoMAEv2-Base, 100-step proprioception and 224-step audio/IMU/pressure go through Transformers, fingertip features are fused across modalities, time and fingers by weight-shared factorized attention, and everything is aggregated by attention fusion into a post-lift stability probability. Under 5-fold object-disjoint cross-validation, vision+proprioception+touch reaches 83.55% and removing touch costs 3.99 points (79.56%). Tactile spatial resolution matters monotonically (79.40% at 1x1 up to 83.55% at 112x112), and freezing touch to its final frame costs 2.07 points, so the useful signal is high-resolution and dynamic. Deployed as an online lift-or-regrasp gate (threshold 0.95, up to five regrasps) on 20 objects excluded from training, the visuo-tactile gate reaches 82.0% success among executed lifts, 10.5 points above a non-tactile gate. The ~1.2 TB dataset is public on OPARA.
Paper Metadata
Title: Temporal Visuo-Tactile Learning for Dexterous Grasp Stability
Authors: Ken Nakahara, Aleksei Buvailik, Prokhor Kotov, Roberto Calandra (LASR Lab, TU Dresden, Germany)
Links: arXiv:2610.10283 (submitted 2026-10-07) · project page lasr-lab.github.io/dexterous-grasp-stability · dataset on OPARA (~1.2 TB)
Code status: released. The repository lasr-lab/dexterous-grasp-stability (MIT for code, CC BY-NC-ND 4.0 for data) ships an HDF5 trial browser (data_viewer.py) and a multimodal resampler (resample.py); model training code is not included.
One-Sentence Summary
Using four Digit 360 fingertips and 10,000 real multi-fingered grasps, this paper shows that the answer to "is this grasp stable before we lift?" lives mostly in high-resolution, time-varying touch, and that turning this judgment into an online gate raises real-robot success among executed lifts by 10.5 percentage points over a non-tactile gate.
Background and Motivation
Stable grasping with a multi-fingered hand is not a pose-selection problem. Stability depends on how contacts form, how contact forces evolve while the fingers close, how load is shared across fingertips, and whether those contacts still hold when the hand starts to lift. Almost all of this happens inside a few square centimeters of interface between fingertip and object. An external camera sees the back of the hand and the silhouette of the object; it does not see the contact. Human manipulation works the other way around: the classic work of Johansson and Flanagan shows that people rely on fingertip tactile signals throughout grasping and manipulation, and everyday human grasp success is close to perfect.
Much of the robotic grasping literature took a different route: estimate a grasp pose from visual observations and execute it with a two-finger parallel gripper. That route works well enough for bin picking because the parallel jaw outsources stability to geometric interlock and force control, sidestepping the question of predicting the outcome before lifting. With an articulated multi-fingered hand, contact configurations become diverse, enveloping grasps become normal, and the margin of vision-only pose selection shrinks quickly.
The visuo-tactile learning thread has existed for a while. Calandra et al. showed in "The Feeling of Success" (2017) that touch helps predict grasp outcomes end to end, and the 2018 follow-up used the learned predictor for model-guided regrasping. Those results, however, were obtained with parallel grippers and early GelSight/DIGIT sensors. They cannot answer the multi-fingered questions: which sensing modalities carry the most information about stability, how much do tactile spatial detail and temporal dynamics each contribute, and how should these signals be modeled?
The hardware changed recently. Digit 360 packs an internal fisheye camera, a microphone, a barometer and an IMU into a hemispherical fingertip, so "touch" is no longer a scalar but a bundle of heterogeneous time series; representation-learning work such as Sparsh-X pretrained multisensory touch encoders on roughly one million contact-rich interactions. Once sensors are rich enough and data cheap enough, a systematic answer becomes possible for the first time. This paper spends that opportunity fully: it does not propose a flashy new module, it quantifies three engineering questions one by one with 10,000 grasps - what to sense, how much spatial and temporal detail to keep, and how to fuse it.
The contributions are therefore threefold. First, a large public multimodal dataset of 10,000 multi-fingered grasp trials over 200 objects, with synchronized external RGB-D, proprioception and four fingertip tactile streams recorded throughout each grasp. Second, a systematic comparison of modality combinations and static versus temporal encoders under object-disjoint cross-validation, plus controlled input ablations that isolate the roles of tactile spatial resolution and temporal dynamics. Third, deployment of the learned predictor as an online lift-or-regrasp gate on the real robot, achieving higher success among executed lifts on unseen objects than a non-tactile gate.
Preliminaries
Vision-based tactile sensors. These sensors convert deformation of a compliant contact surface into images: GelSight recovers contact geometry with photometric stereo, TacTip uses biomimetic soft optical fingertips, and DIGIT made the design compact and cheap enough to mount on fingers. Digit 360 is the newest generation of this lineage, combining a hemispherical contact surface with an internal fisheye camera plus audio, pressure and inertial channels.
Grasp stability prediction. The task is to predict, from observations taken before lifting, whether the object will be lifted and held without dropping. Analytical grasp quality measures and force-closure metrics require known contact geometry and friction models that are hard to obtain reliably on real hardware; hand-tuned tactile controllers with event thresholds do not cover diverse objects. End-to-end learning turns the problem into binary classification and lets the data express contact mechanics implicitly.
The temporal multimodal toolbox. Pretrained video models (VideoMAEv2) provide strong priors for short clips; factorized attention processes the time and finger dimensions separately instead of paying for joint attention over their product; modality masks let the network run forward when a sensor stream is missing. These three tools form the skeleton of the network in this paper.
Method in Detail
Task formulation. Given a sensory observation window $\mathbf{s}_{t:t+T}$ from the pre-lift grasp phase, the goal is to learn a function that predicts the probability of a successful grasp,
$$f_{\theta}(\mathbf{s}_{t:t+T}) = p(o=1 \mid \mathbf{s}_{t:t+T}) \in [0,1]$$
where the label $o=1$ means the object was lifted and held without dropping. The network is trained on the 10,000-trial dataset with binary cross-entropy,
$$\mathcal{L}(\theta) = -\frac{1}{N}\sum_{i=1}^{N}\left[\, o_i \log f_{\theta}(\mathbf{s}_i) + (1-o_i)\log\left(1-f_{\theta}(\mathbf{s}_i)\right) \right]$$
The decisive constraint in this formulation is temporal direction: inputs may only come from the window before lift onset. This is not a modeling preference but a deployment requirement, because a gate that decides after the lift has started cannot trigger a regrasp.
Observation window and input representations. The window is three seconds long, anchored at grasp initiation and ending before lift onset. External RGB and the four tactile cameras contribute 16-frame clips, resized to $256\times 256$ and cropped to $224\times 224$ with crop coordinates consistent across frames of a clip; training uses one center-cropped and one randomly cropped version of each sample, doubling samples per epoch, while evaluation uses center crops. Proprioception is a 100-step sequence of 16 hand joint angles plus a 6-D end-effector pose. Per fingertip, audio, IMU and pressure become 224-step sequences: 256-bin log-magnitude spectra, three-axis acceleration, and scalar pressure readings. Image branches use pretrained VideoMAEv2-Base followed by temporal Transformers, with the tactile encoder shared across fingers; the remaining streams each get a Transformer sequence encoder.
Intra-finger cross-modal fusion. Inside the touch branch, encoded camera, audio, IMU and pressure features are resampled onto a common temporal grid and merged at each time step by a modality-fusion Transformer whose weights are shared across fingers. Because real sensors can drop out entirely, both touch fusion and cross-modal aggregation use modality masks. In the released resample.py, the mask convention in _fill_finger_touch is that a channel is 1.0 where the modality is missing, so an all-ones row means that fingertip contributed nothing.
Factorized attention over time and fingers. The four-finger temporal tokens then pass through weight-shared factorized Transformer blocks that attend along time and along fingers separately, rather than over the joint time-by-finger sequence. Lightweight finger-specific residual adapters absorb per-finger systematic bias (assembly differences, gel aging, calibration drift), and temporal pooling reduces each fingertip to a single token. The design intent is explicit: cross-finger commonality is learned by shared weights, finger individuality by small residuals, instead of training four separate parameter sets.
Cross-modal attention fusion and prediction head. The RGB token, the proprioceptive token, the four fingertip tokens and 16 learnable fusion tokens are concatenated in a shared 512-dimensional space and processed jointly by three Transformer layers with eight attention heads. The learnable tokens act as scratch space where visual evidence about object pose and tactile evidence about contact deformation can reference each other. Mask-aware pooling over the updated branch tokens feeds a two-layer MLP head (512-dimensional hidden layer, LayerNorm, GELU, dropout 0.2) followed by a sigmoid that outputs the grasp success probability.
Online gate: lift or regrasp. At deployment the robot reaches the vicinity of the object and closes the hand toward a sampled grasp configuration; the predictor evaluates the three-second window anchored at grasp initiation; the robot lifts only when the predicted probability exceeds a predetermined threshold, otherwise it opens the hand, applies a small relative displacement to the end-effector, and samples another grasp. In the real-robot experiments the threshold is fixed at 0.95 (set empirically before testing); each episode allows one initial grasp plus up to five regrasps, and six rejections trigger a no-lift reset that returns the robot home and restores the object's resting configuration. The gate is therefore not a controller but a trial-budget allocator: it decides which grasps deserve to spend a lift.
Data and labels: turning annotation into a physical measurement. At the start of each trial the object is placed randomly in a $200\,\mathrm{mm}\times 400\,\mathrm{mm}$ tabletop workspace ($x\in[650,850]\,\mathrm{mm}$, $y\in[-200,200]\,\mathrm{mm}$ in the grasping-arm base frame). SAM 3 segments the object in the external RGB image from a text prompt, and the object center in the arm base frame is estimated from the mask centroid, the depth image and calibrated camera extrinsics. The end-effector target (defined at the base of the middle finger) receives translational perturbations (default $[-50,50]\,\mathrm{mm}$ in $x$ and $y$, $[-25,25]\,\mathrm{mm}$ in $z$, adjusted for object size and dataset balance) and a yaw sampled from $[-20^{\circ},20^{\circ}]$. The 16 hand joint targets are sampled independently and uniformly from calibrated ranges (15-20 degrees for most joints, 45 degrees for the index, middle and ring MCP joints). The hand closes through 50 linearly interpolated position-control waypoints at a nominal 50 Hz. After closure and the observation window, the robot lifts by 100 mm; the trial is labeled $o=1$ if the object is lifted and held for 3 s without dropping. The automatic labeling rule is a height test,
$$o = 1 \iff h_{\text{post}} - h_{\text{pre}} \geq 50\,\mathrm{mm}$$
with height estimated from the SAM 3 mask, depth and extrinsics; when post-lift detection fails (for example under hand occlusion) the label falls back to a logistic model over fingertip displacements inferred from the difference between commanded and observed 16-DoF hand configurations. Automatic labels agree with manually verified labels on 92.34% of all 10,000 trials (9234/10000); the confusion matrix is 4361 stable-stable, 415 automatic-stable but manual-unstable, 351 automatic-unstable but manual-stable, and 4873 unstable-unstable. The dataset is nearly balanced: 47.1% stable, 52.9% unstable.
Pipeline diagram drawn from the actual architecture of this paper:
flowchart LR
subgraph WIN["3s pre-lift window (anchored at grasp initiation)"]
RGB["external RGB 16 frames"]
PRO["proprio 100 steps
16 joint angles + 6D EE pose"]
TC["tactile camera 16 frames x4"]
AU["audio 224 steps x4
256-bin log spectra"]
IM["IMU 3-axis acc x4"]
PR["pressure scalar x4"]
end
RGB --> VE["VideoMAEv2-Base + temporal Transformer"]
PRO --> PT["Transformer sequence encoder"]
TC --> TE["VideoMAEv2-Base + temporal Transformer
weights shared across fingers"]
AU --> AT["Transformer"]
IM --> IT["Transformer"]
PR --> RT["Transformer"]
TE --> IF["intra-finger modality-fusion Transformer
common temporal grid + modality mask"]
AT --> IF
IT --> IF
RT --> IF
IF --> FA["factorized attention over time and fingers
+ finger-specific residual adapters"]
FA --> POOL["pool over time -> one token per fingertip"]
VE --> FUS["attention fusion: 6 sensing tokens + 16 learnable tokens
3 layers x 8 heads, dim 512"]
PT --> FUS
POOL --> FUS
FUS --> MP["mask-aware pooling"]
MP --> HEAD["2-layer MLP head + sigmoid"]
HEAD --> GATE{"p >= 0.95 ?"}
GATE -->|yes| LIFT["lift 100 mm"]
GATE -->|no| REGRASP["release, small EE offset,
resample grasp (max 5 regrasps)"]
The diagram mirrors Fig. 3 of the paper: four sensing branches merge within each finger, then across fingers, and finally with vision and proprioception in global attention; the 0.95 threshold and the five-regrasp budget come from the deployment protocol in Sec. VI-B.
Experimental Results
Protocol. Offline evaluation uses 5-fold object-disjoint cross-validation: each fold holds out all trials of a subset of objects, so accuracy measures generalization to unseen object instances. Every model trains for 100 epochs with AdamW, batch size 4, initial learning rate $10^{-6}$, weight decay $10^{-4}$, cosine schedule with a 0.03 warmup ratio; accuracy is computed from the final-epoch checkpoint at a 0.5 threshold and reported as mean and standard deviation over the five folds. Each input ablation trains a separate model with the modification applied in both training and evaluation.
Modalities: touch is the main increment, and the tactile camera is the main part of touch. In the VideoMAEv2-based modality ablation, touch-only, vision-touch and vision-proprioception-touch reach mean accuracies of 82.64%, 83.52% and 83.55%. Removing touch from the full model costs 3.99 percentage points, dropping to 79.56%. The full model beats touch-only by only 0.91 points (paired t-test across five matched folds, $p=0.035$) and nearly matches vision-touch, indicating limited additional benefit from proprioception once vision is present. Restricting the touch branch to the tactile camera yields 83.41%, essentially the full-touch result of 83.55%; using audio, IMU or pressure as the sole tactile modality gives 79.60-79.94%, on par with the non-tactile baseline of 79.56%.
| Touch-branch inputs (vision + proprioception retained) | Accuracy [%] |
|---|---|
| All Digit 360 modalities | 83.55 ± 1.58 |
| Tactile camera only | 83.41 ± 1.30 |
| Audio only | 79.60 ± 1.46 |
| IMU only | 79.94 ± 1.50 |
| Pressure only | 79.93 ± 1.10 |
| No tactile streams (vision + proprioception) | 79.56 ± 1.60 |
Backbones: temporal modeling matters more than the pretraining source. Among three encoding designs, the static ResNet-50+MLPs baseline that sees only the final observation reaches 79.79%, showing that the final pre-lift state is already informative; ResNet-Transformer rises to 81.42%; the VideoMAEv2-based model reaches 83.55%. Replacing raw tactile RGB with absolute differences from each clip's first frame (background subtraction) gives 82.58%, slightly below raw frames at 83.55%, consistent with VideoMAEv2's pretraining on large-scale RGB video: raw frames keep color, illumination and stable gel-appearance cues. Pooling the four fingertip representations into a single touch token before cross-modal fusion (local fusion) gives 83.24% versus 83.55% for global fusion, a 0.31-point gap that suggests most cross-finger context is already encoded by the preceding factorized attention.
| Encoding design (vision + proprioception + touch) | Input window | Accuracy [%] |
|---|---|---|
| Static ResNet-50 + MLPs | final frame only | 79.79 |
| ResNet-50 + temporal Transformer | 3 s window | 81.42 |
| VideoMAEv2-Base + temporal Transformer | 3 s window | 83.55 ± 1.58 |
| same, background-subtracted tactile frames | 3 s window | 82.58 ± 1.94 |
| same, four fingers pooled to one token early | 3 s window | 83.24 ± 1.61 |
Spatial resolution and temporal dynamics: two controlled ablations give causal-style evidence. Because the backbone comparison also changes architecture, capacity and pretraining, the paper isolates time by perturbing inputs and space by area downsampling. Spatially, downsampling tactile frames to $1\times1$ (keeping only per-frame mean RGB) yields 79.40%, indistinguishable from the non-tactile baseline of 79.56%; accuracy climbs steeply to 82.51% at $28\times28$, then more gradually to 83.55% at both $112\times112$ and $224\times224$, so spatially resolved contact patterns matter with diminishing returns at the top. Temporally, repeating a stream's final observation throughout the window (keeping the final state, removing within-window variation) lowers vision-only and touch-only accuracy while leaving proprioception-only nearly unchanged; in the full model, repeating touch costs 2.07 points and additionally repeating vision or proprioception costs more; applying a common random permutation to the external RGB and the four tactile-camera clips (preserving all frames and their mutual alignment, destroying order) costs 1.62 points. The conclusion is that the useful signal is not only what contact looks like but how it changes.
| Tactile frame spatial resolution | 1×1 | 28×28 | 112×112 | 224×224 |
|---|---|---|---|---|
| Mean accuracy [%] | 79.40 | 82.51 | 83.55 | 83.55 |
Fingers: thumb opposition carries information vision cannot recover. Restricting finger-specific inputs (joint angles, plus the corresponding Digit 360 streams when touch is enabled) to finger subsets, accuracy across subsets spans 11.84 points with proprioception alone, 0.59 points with vision plus proprioception, and 3.82 points with vision, proprioception and touch. External vision therefore recovers almost all of the removed joint state but not the removed fingertip contact. Relative to vision-proprioception with the same subset, touch adds 3.06 points for the thumb, 1.35 for the middle finger and 0.14 for the index; the all-finger configuration is best (83.55%), closely followed by thumb-middle (83.41%). The descriptive analysis tells the same story: grouping trials by the number of detected contacts in the final pre-lift frame, success rises monotonically from 6.9% at zero contacts to 26.5%, 58.2%, 70.0% and 76.8% at one through four contacts; among two-contact sets, those containing the thumb reach 47.3-66.4% versus 19.6-25.2% without it. A heuristic rule built on the same detector - stable when the thumb and at least one other fingertip are in contact, excluding geometric finger-finger self-contact - reaches 73.07% accuracy over all 10,000 trials, far above the 52.9% majority-class baseline but below every learned touch-only model (78.59% static, 79.98% ResNet-Transformer, 82.64% VideoMAEv2). Binary contact detection captures part of the stability signal; the rest lives in the spatial pattern within each fingertip, the relations across fingers, and their evolution over time.
Real robot: both the gain and the cost of gating are reported honestly. Four policies were evaluated on 20 objects excluded from the training set, with 10 executed lifts per object: the tactile gate (vision+proprioception+touch) reaches 82.0% (164/200) success among executed lifts, versus 71.5% for the non-tactile gate, 47.5% for the random gate and 49.5% for direct lifting. Object-level paired t-tests give $p=0.003$ for tactile versus non-tactile and $p<0.001$ for tactile versus random, while random versus direct lifting shows no detectable difference ($p=0.48$), so the improvement is not explained by abstention and retry alone. Counting abstentions in the denominator, conservative episode-level success rates are 59.2%, 54.8%, 38.2% and 49.5%, preserving the same ordering. A clean discrimination probe comes from the random gate, whose acceptances are independent of the predictor score: among its 200 executed lifts, those the predictor scored above threshold succeeded 86.5% of the time (77/89) while those below threshold succeeded 16.2% (18/111). The predictor discriminates stable from unstable grasps rather than winning by lifting less. The four policies consumed 779, 690, 764 and 200 grasp candidates with 77, 61, 49 and 0 no-lift resets: the tactile gate spends more grasp attempts to buy fewer failed lifts.
| Real-robot policy (20 unseen objects) | Success among executed lifts | Episode-level (abstentions counted) | Grasp candidates | No-lift resets |
|---|---|---|---|---|
| Tactile gate (V+P+T) | 82.0% (164/200) | 59.2% | 779 | 77 |
| Non-tactile gate (V+P) | 71.5% | 54.8% | 690 | 61 |
| Random gate (accept 25%) | 47.5% | 38.2% | 764 | 49 |
| Direct lift | 49.5% | 49.5% | 200 | 0 |
Figure 7 (paper Fig. 11): per-object and mean success among executed lifts for the four policies on 20 unseen objects. The visuo-tactile gate (red) leads on most objects, 82.0% mean versus 71.5% for the non-tactile gate.
Figure 6 (paper Fig. 6): modality and encoding-backbone ablations under 5-fold object-disjoint CV. Touch-enabled models lead, multi-frame models beat the static single-frame baseline, and the VideoMAEv2-based model is best.
Figure 3 (paper Fig. 3): architecture. Vision and proprioception each get one branch; the four fingertip tactile streams are fused within each finger across modalities, then across time and fingers by factorized attention with finger-specific adapters, before global attention fusion produces the grasp success probability.
Figure 1 (paper Fig. 1): overview. During grasp execution the robot observes the object and fingertip contacts through external vision, proprioception and four Digit 360 sensors; the temporal predictor estimates post-lift stability and gates lift-or-regrasp decisions online.
Figure 5 (paper Fig. 5): automated collection timeline. Localization, randomized reaching and grasping, lifting, and stability labeling from object height changes, with continuous multimodal recording.
Dataset scale and label quality. Subsampling experiments show accuracy generally increasing with data; among single-modality models vision-only degrades fastest as data shrinks, while touch-enabled models stay comparatively robust to reduced object coverage. On labels, the automatic pipeline agrees with manual verification on 92.34% of trials; in follow-up work on the same recordings a learned post-lift detector reaches 99.14% agreement (fusing proprioception, tactile video, contact audio and external vision), 98.60% using only hand-local signals, and 97.75% with proprioception alone, which means accurate automatic labeling need not depend on an external camera view or its calibration. The object set itself is worth recording: 200 objects spanning 1.94-244.4 g in mass (median 32.34 g) and 42.3-1210 cm3 in bounding-box volume (median 216.3 cm3), with 105 rigid and 95 deformable objects and plastic (107), silicone (23) and metal (16) as the largest material groups; per-object success rates range from 7.1% (a Meta Quest controller) to 90.6%, with 17 objects below 25% and 15 above 75%.
Limitations
Stated by the authors: mass is not an input. On the three heaviest objects (a 244 g Nutella jar, a 229 g tuna can and similar), real-robot success stayed at 40-60%. The authors' explanation is that locally plausible pre-lift contacts may not ensure sufficient support under load, because object mass is not an explicit predictor input. In other words the predictor learns stability of contact geometry and dynamics, not the ratio of contact forces to weight.
Stated by the authors: it is open whether non-camera tactile streams are uninformative or merely underused. Audio, IMU and pressure used alone sit close to the non-tactile baseline. The authors explicitly leave open whether these signals carry little information in this task or whether the shared encoder and fusion design fails to exploit them; encoders such as Sparsh-X, pretrained with modality-specific temporal windows, were excluded because the predictor consumes a synchronized three-second window, and the authors list them as a possible improvement.
Stated by the authors: the contact analyses are observational. In the contact-count and contact-pattern analyses, contacts are detected rather than controlled, so each comparison groups trials that differ in many other ways; the detector thresholds were selected by visual inspection, and rare contact sets rest on few trials with wide confidence intervals.
Our assessment: a single operating point and an abstention-aware denominator. The 0.95 threshold is one empirically fixed operating point chosen before testing; the paper reports no threshold sweep or precision-recall curve. Success among executed lifts excludes no-lift resets from the denominator; under the conservative episode-level metric the tactile gate is at 59.2% versus 49.5% for direct lifting, and it pays 77 resets to get there. Both denominators should be read together, otherwise the gate looks cheaper than it is.
Our assessment: open-loop grasping plus a gate is not closed-loop tactile control. Hand closure is open-loop position control through 50 waypoints; touch neither corrects the trajectory nor modulates grip force. The gate only filters which grasps get to lift. The 10.5-point gain should therefore be read as better allocation of trial budget, not as better grasp control. The authors themselves list closed-loop tactile control during lifting and action-conditioned predictors as future work.
Conclusion and Outlook
The value of this paper is not a new module but a set of auditable measurements on a question that had been answered by intuition: how much does touch actually help dexterous grasping, and where? The modality ablation gives the marginal value of touch (+3.99 points), the resolution ablation the marginal value of spatial detail (+4.15 points from 1x1 to 112x112), the temporal ablation the marginal value of dynamics (-2.07 points when tactile time is frozen), the finger ablation the part of multi-finger contact that vision cannot recover (+3.06 points for the thumb), and the real-robot gate the net deployment gain (+10.5 points). Five evidence chains point to the same conclusion: high-resolution, time-varying fingertip contact is the strongest learnable signal for multi-fingered grasp stability.
The authors' stated directions include feeding candidate actions as inputs and continually learning from regrasp trials collected during deployment to grow the predictor into an action-conditioned model that scores reaching and hand-closure trajectories; using the gate as a proposal-agnostic pre-lift critic for learned grasp planners; guiding closed-loop tactile control during lifting; adding multi-view or depth geometry; and extending binary labels to finer events such as slip onset or contact loss. For teams building dexterous manipulation systems the paper also leaves a blunt engineering conclusion: before buying a more expensive hand, put high-resolution touch on the fingertips and let a model that has seen ten thousand failures decide whether to lift.
Golden Quote
"Manipulating objects around us almost always starts with grasping them." - and before the lift, the fingertips already know the answer.



