Inference-time Policy Steering via Vision and Touch
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
TacSushi is a tactile‑grounded, Cosmos3‑based world‑action policy for dexterous sushi manipulation. It encodes RGB, language, and hand state, fusing fingertip tactile data via feature‑wise gated fusion, and learns from future‑consequence predictions while excluding failed actions from imitation. Trained on 340 successful and 50 failed trials, TacSushi achieves 68.3% in‑distribution and 37.5% out‑of‑distribution success, outperforming baselines that lack future‑consequence supervision or use direct tactile concatenation.
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions.
arXiv:2609.09119v1 Announce Type: cross Abstract: Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision...
DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
TacForcing is a streaming action‑generation framework that incorporates execution‑time tactile feedback into contact‑rich manipulation. It replaces a separate reactive controller with a streaming action expert that conditions actions on evolving tactile observations, and introduces Execution‑Aware Tactile Attention (EATA) to focus tactile conditioning on actions near execution. The method achieves 65% success in six simulated UniVTAC tasks and 69% in three real‑world contact‑rich manipulation tasks, outperforming strong baselines.
Accurate contact prediction is useful for robotic manipulation only if it supports effective decisions. We investigate this connection using a compact, randomly initialized visuotactile world model, t...
arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.
arXiv:2606. 04708v1 Announce Type: cross Abstract: Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging.