arXiv AI By Tatsuya Kamijo, Mai Nishimura, Nodoka Shibasaki, Jeremy Siburian, Cristian C. Beltran-Hernandez, Masashi Hamaya

Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist

Read the original on arXiv AI →

The paper introduces TaMeSo‑bot, a soft‑wrist robotic system that uses tactile memory to perform robust object insertion tasks. It employs a Masked Tactile Trajectory Transformer (MAT³) to jointly model actions, tactile cues, force‑torque data, and proprioception, learning spatiotemporal representations through masked token prediction. Experiments on peg‑in‑hole tasks show that MAT³ outperforms baselines and adapts well to unseen pegs and conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 18

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.

By Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, Renjing Xu
arXiv AI
Sep 18

TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation

TacSushi is a tactile‑grounded, Cosmos3‑based world‑action policy for dexterous sushi manipulation. It encodes RGB, language, and hand state, fusing fingertip tactile data via feature‑wise gated fusion, and learns from future‑consequence predictions while excluding failed actions from imitation. Trained on 340 successful and 50 failed trials, TacSushi achieves 68.3% in‑distribution and 37.5% out‑of‑distribution success, outperforming baselines that lack future‑consequence supervision or use direct tactile concatenation.

By Haodi Hu, Kaen Kogashi, Toshiaki Koike-Akino