The paper presents a 652,157‑parameter action‑conditioned visuotactile world model designed for lifting tasks, integrating behavior cloning, policy learning in imagination, reactive implicit Q‑learning, and model‑assisted force feedback. Experiments on 120 fresh MuJoCo environments and additional ID environments show that visuotactile dynamics reduce force‑action‑effect mean absolute error from 0.413 N to 0.338 N, and model‑assisted feedback boosts force‑budgeted success from 73.3 % to 93.3 %. Imagined reinforcement learning achieves 11.9 % pooled joint success compared to 25.0 % for reactive IQL, with further stress testing adding 330 executions.
By Qinzhen Ma (Rice University)
TacSushi is a tactile‑grounded, Cosmos3‑based world‑action policy for dexterous sushi manipulation. It encodes RGB, language, and hand state, fusing fingertip tactile data via feature‑wise gated fusion, and learns from future‑consequence predictions while excluding failed actions from imitation. Trained on 340 successful and 50 failed trials, TacSushi achieves 68.3% in‑distribution and 37.5% out‑of‑distribution success, outperforming baselines that lack future‑consequence supervision or use direct tactile concatenation.
By Haodi Hu, Kaen Kogashi, Toshiaki Koike-Akino
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
By Siyu Ma, Yuqi Liang, Chang Yu, Yunuo Chen, Hao Su, Yixin Zhu, Yin Yang, Chenfanfu Jiang
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
By Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang
arXiv:2609.15726v1 Announce Type: cross
Abstract: Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not conve...
By Zhenjie Yang, Yideng Zhang, Dongjie Zhang, Chenyu Jiang, Xianshuai Liu, Yufeng Li, Zuhao Ge, Xingyu Jiao, Zheng Zhang, Kaiyu He, He Wang, Yuwen Zhong, Yi Deng, Muyun Jiang, Xianliang Huang, Haisheng Su, Donghang Zhang, Jian Zhang, Xue Yang, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
By Yilin Wu, Zilin Si, Zeynep Temel, Oliver Kroemer, Andrea Bajcsy
DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.
By Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, Renjing Xu
arXiv:2607. 09218v1 Announce Type: cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
By Rishabh Madan, Angchen Xie, Samantha Saak, Andres Blanco, Dohyeok Lee, Sarah Grace Brown, Yunting Yan, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee
arXiv:2607. 09218v2 Announce Type: replace-cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
By Rishabh Madan, Angchen Xie, Samantha Saak, Andres Blanco, Dohyeok Lee, Sarah Grace Brown, Yunting Yan, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee
arXiv:2609.13235v1 Announce Type: cross
Abstract: Networked manipulation endpoints couple perception to actuation across compute- and bandwidth-limited links, yet commonly exchange dense geometric st...
By Md Selim Sarowar, Sungho Kim
arXiv:2604. 25897v2 Announce Type: replace-cross Abstract: Contact variability, sensing uncertainty, and external disturbances make grasp execution stochastic.
By Clinton Enwerem, Shreya Kalyanaraman, John S. Baras, Calin Belta
arXiv:2609.16697v1 Announce Type: cross
Abstract: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interv...
By Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen, Peilin Zhao, Zongqing Lu, Gao Huang, Steven Hoi, Dacheng Tao, Deheng Ye