DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.
arXiv:2609.09119v1 Announce Type: cross
Abstract: Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision...
By Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
By Siyu Ma, Yuqi Liang, Chang Yu, Yunuo Chen, Hao Su, Yixin Zhu, Yin Yang, Chenfanfu Jiang
TacSushi is a tactile‑grounded, Cosmos3‑based world‑action policy for dexterous sushi manipulation. It encodes RGB, language, and hand state, fusing fingertip tactile data via feature‑wise gated fusion, and learns from future‑consequence predictions while excluding failed actions from imitation. Trained on 340 successful and 50 failed trials, TacSushi achieves 68.3% in‑distribution and 37.5% out‑of‑distribution success, outperforming baselines that lack future‑consequence supervision or use direct tactile concatenation.
By Haodi Hu, Kaen Kogashi, Toshiaki Koike-Akino
arXiv:2609.24976v1 Announce Type: cross
Abstract: Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple pre...
By Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan
arXiv:2609.34182v2 Announce Type: replace-cross
Abstract: Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleope...
By Wenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu, Tengyu Liu, Kaifeng Zhang, Chuan Wen, Siyuan Huang
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
By Yilin Wu, Zilin Si, Zeynep Temel, Oliver Kroemer, Andrea Bajcsy
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
By Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang
VT-MUSE is a multimodal unified sequential representation learning framework for visuotactile manipulation. It uses a two‑stage approach: first, modality‑specific encoders are jointly adapted with cross‑modal temporal alignment and masked‑view consistency; second, a conditional variational latent model processes masked visual sequences and full tactile histories, with auxiliary decoders reconstructing recent visual observations and predicting tactile depth changes. The resulting representation is fed into a lightweight Transformer policy via gated cross‑attention, achieving an 11‑percentage‑point improvement over the strongest baseline in simulation and significant gains in real‑world experiments.
By Congsheng Xu, Qiaochu Yang, Fangyuan Shi, Yifan Han, Baijun Chen, Yiming Wang, Haonan Zhao, Daolin Ma, Xiaokang Yang, Hesheng Wang
arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.
By Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han, Ruohan Gao
arXiv:2606. 31694v1 Announce Type: cross Abstract: For robots manipulating open-world objects, tactile representations must generalize to unseen materials.
By Jingbo He, Michael F\"arber, Roberto Calandra
arXiv:2608.29601v2 Announce Type: replace-cross
Abstract: We present $N_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale mul...
By NeoteAI Team, Fudan TEAI Team