arXiv:2608.29601v2 Announce Type: replace-cross
Abstract: We present $N_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale mul...
By NeoteAI Team, Fudan TEAI Team
arXiv:2608.29601v1 Announce Type: cross
Abstract: We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale m...
By NeoteAI Team, Fudan TEAI Team
Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.
By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
arXiv:2609.14783v1 Announce Type: cross
Abstract: Robots need touch to manipulate objects safely and reliably, as many properties, such as softness, texture, and contact stability, are hard to infer...
By Mashood M. Mohsan, Muhayy Ud Din, Binzhao Xu, Ahmad Abubakar, Irfan Hussain
The paper introduces BIFTA, a Brain‑Inspired Few‑Shot Tactile Adaptation framework that enables a frozen encoder to adapt quickly to an unknown tactile sensor using only a small labeled support set. It preserves pretrained representations via dual‑view statistical memory, builds support‑conditioned spectral graphs to correct sensor‑dependent feature neighborhoods, and employs uncertainty‑gated recurrent propagation to reinforce reliable cross‑query evidence. Benchmarks on three tactile datasets demonstrate that BIFTA dramatically improves adaptation performance, achieving an 87.09% mean Sparsh accuracy on SITR with just 10% labeled data—an increase of 47.22 percentage points over the best prior method.
By Boheng Liu, Ziyu Li, Xia Wu
arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.
By Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han, Ruohan Gao
DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.
By Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, Renjing Xu
arXiv:2606. 11767v1 Announce Type: cross Abstract: Blind grasping with a dexterous hand is a crucial manipulation capability.
By Shengcheng Luo, Xiyan Huang, Zhe Xu, Wanlin Li, Ziyuan Jiao, Chenxi Xiao
BIDETA is a gradient‑free framework that adapts pretrained tactile models to new sensors using only a few labeled target contacts. It preserves pretrained representations while repairing sensor‑dependent feature neighborhoods through rapid support memory, support‑conditioned spectral graphs, and reliability‑gated recurrence. Experiments on multiple datasets show that BIDETA dramatically improves accuracy and speeds up adaptation compared to prior methods.
By Boheng Liu, Lan Wei, Ziyu Li, Chenghua Duan, Qing Li, Dandan Zhang, Xia Wu
arXiv:2609.21449v1 Announce Type: new
Abstract: World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future vis...
By Xuancheng Zhang, Xuetao Liu, Qianying Tang, Jizhe Wang, Zhijing Cheng, Bochen Lin, Haoran Wen, Ming Li, Kun Zhan, Yu Liu
ControlTac is a two‑stage framework that generates realistic tactile images conditioned on a single reference image, contact force, and contact pose. By incorporating these physical priors, it produces realistic samples across different sensors and captures task‑relevant variations. Experiments in object insertion, imitation learning, and object weighting show that datasets augmented with ControlTac consistently improve performance in dynamic real‑world settings.
By Dongyu Luo, Kelin Yu, Amir-Hossein Shahidzadeh, Cornelia Ferm\"uller, Yiannis Aloimonos, Ruohan Gao
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
By Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang