Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper presents a 652,157‑parameter action‑conditioned visuotactile world model designed for lifting tasks, integrating behavior cloning, policy learning in imagination, reactive implicit Q‑learning, and model‑assisted force feedback. Experiments on 120 fresh MuJoCo environments and additional ID environments show that visuotactile dynamics reduce force‑action‑effect mean absolute error from 0.413 N to 0.338 N, and model‑assisted feedback boosts force‑budgeted success from 73.3 % to 93.3 %. Imagined reinforcement learning achieves 11.9 % pooled joint success compared to 25.0 % for reactive IQL, with further stress testing adding 330 executions.
TacSushi is a tactile‑grounded, Cosmos3‑based world‑action policy for dexterous sushi manipulation. It encodes RGB, language, and hand state, fusing fingertip tactile data via feature‑wise gated fusion, and learns from future‑consequence predictions while excluding failed actions from imitation. Trained on 340 successful and 50 failed trials, TacSushi achieves 68.3% in‑distribution and 37.5% out‑of‑distribution success, outperforming baselines that lack future‑consequence supervision or use direct tactile concatenation.
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
arXiv:2609.15726v1 Announce Type: cross Abstract: Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not conve...
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.