Inference-time Policy Steering via Vision and Touch
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
arXiv:2607. 09218v2 Announce Type: replace-cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
arXiv:2607. 09218v1 Announce Type: cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
arXiv:2609.09119v1 Announce Type: cross Abstract: Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision...
DeCAL is a vision‑language‑action model designed for dexterous manipulation that incorporates tactile sensing through adaptive visuo‑tactile fusion and latent co‑imagination. It uses a Mixture‑of‑Transformers architecture with specialized experts for understanding, imagination, and action, enabling efficient information flow and dynamic regulation of tactile inputs. Experiments show DeCAL achieves state‑of‑the‑art performance, with a 71% average success rate and 83.4% progress success rate, and generalizes well to unseen scenarios.
DexTouch-WM is an action‑conditioned world model that learns from scalable human touch to predict future RGB observations and bilateral tactile dynamics for dexterous robot manipulation. By using compatible piezoresistive arrays on both human and robot hands and retargeting human motion into the robot action space, the model can be supervised with human interaction data while keeping a fixed amount of real‑robot supervision. Experiments show that adding up to 100 hours of human interaction improves robot‑domain visual, geometric, and contact prediction, and the model can serve as a surrogate environment for policy evaluation and synthetic trajectory generation.
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.
arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.
TacForcing is a streaming action‑generation framework that incorporates execution‑time tactile feedback into contact‑rich manipulation. It replaces a separate reactive controller with a streaming action expert that conditions actions on evolving tactile observations, and introduces Execution‑Aware Tactile Attention (EATA) to focus tactile conditioning on actions near execution. The method achieves 65% success in six simulated UniVTAC tasks and 69% in three real‑world contact‑rich manipulation tasks, outperforming strong baselines.
arXiv:2608.22067v1 Announce Type: cross Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot action...
DELE-w0.5 is a robotic manipulation framework that predicts future latent states instead of generating full video sequences, thereby inferring robot actions directly from these compact representations. By focusing on physical state changes rather than visual transitions, it reduces model complexity and inference latency. In 480 real‑robot trials across four long‑horizon tasks, DELE‑w0.5 achieved 62.5 % overall task success and 81.3 % macro ordered‑stage progress, outperforming the strongest baseline by 47.5 and 30.7 percentage points.