arXiv Machine Learning

Making Sense of Touch from the Child's View for Contrastive Learning

arXiv:2606. 31943v1 Announce Type: new Abstract: Is the sense of touch a mechanism for human babies' learning of visual concepts?

arXiv AI
Sep 18

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces from video. It uses 500 hours of pressure‑glove recordings and hand‑object interaction data, and introduces TwinTouch‑20H, a dataset of 20 hours of paired visual data where generative models render gloved recordings as bare‑hand observations while preserving tactile labels. The system outperforms prior methods on OakInk2, generalizes qualitatively to natural bare‑hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales.

By Danyan Zhou, Jinxuan Lu, Jiawei Lin, Tianxing Chen, Chuqiao Lyu, Wenbo Ding
Hugging Face Trending Papers
Sep 17

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

TouchSight is a monocular egocentric vision framework that predicts dense full-hand contact forces without tactile sensors. It uses 500 hours of pressure-glove data and a 20-hour TwinTouch-20H dataset where generative models render gloved recordings as bare-hand videos, bridging the appearance gap. The system outperforms previous methods on OakInk2, generalizes to unseen natural bare-hand egocentric videos, and improves as glove supervision increases.

arXiv Computer Vision
Aug 24

VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

VT-MUSE is a multimodal unified sequential representation learning framework for visuotactile manipulation. It uses a two‑stage approach: first, modality‑specific encoders are jointly adapted with cross‑modal temporal alignment and masked‑view consistency; second, a conditional variational latent model processes masked visual sequences and full tactile histories, with auxiliary decoders reconstructing recent visual observations and predicting tactile depth changes. The resulting representation is fed into a lightweight Transformer policy via gated cross‑attention, achieving an 11‑percentage‑point improvement over the strongest baseline in simulation and significant gains in real‑world experiments.

By Congsheng Xu, Qiaochu Yang, Fangyuan Shi, Yifan Han, Baijun Chen, Yiming Wang, Haonan Zhao, Daolin Ma, Xiaokang Yang, Hesheng Wang