arXiv AI By Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

Read the original on arXiv AI →

The paper introduces VISTA, a workspace-level equivariant visuotactile diffusion policy designed for data-efficient imitation learning in contact-rich manipulation tasks. VISTA transforms visual and tactile inputs into spherical tokens, fuses them through permutation-equivariant spherical fusion, and aligns the fused representation with the end-effector orientation to produce spatially consistent actions. Experiments in simulation and real-world robotics demonstrate that VISTA significantly improves data efficiency compared to strong visuotactile imitation learning baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 7

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.

By Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han, Ruohan Gao
arXiv Machine Learning
Jun 11

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.

By Siyu Ma, Yuqi Liang, Chang Yu, Yunuo Chen, Hao Su, Yixin Zhu, Yin Yang, Chenfanfu Jiang
Hugging Face Trending Papers
Sep 8

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.