DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.
arXiv:2609.09119v1 Announce Type: cross Abstract: Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision...
arXiv:2609.21449v1 Announce Type: new Abstract: World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future vis...
DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.