NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation
arXiv:2606. 24450v1 Announce Type: cross Abstract: Perceiving physical contact is fundamental to dexterous manipulation.
arXiv:2606. 11767v1 Announce Type: cross Abstract: Blind grasping with a dexterous hand is a crucial manipulation capability.
arXiv:2606. 24450v1 Announce Type: cross Abstract: Perceiving physical contact is fundamental to dexterous manipulation.
arXiv:2606. 11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks.
arXiv:2607. 09218v2 Announce Type: replace-cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
arXiv:2607. 09218v1 Announce Type: cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break.
arXiv:2607. 03723v1 Announce Type: cross Abstract: Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry.
arXiv:2606. 24712v1 Announce Type: cross Abstract: Humans effortlessly locate and identify objects by touch alone, even without vision.
arXiv:2606. 14981v1 Announce Type: cross Abstract: Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution.
arXiv:2606. 10683v1 Announce Type: cross Abstract: Dexterous hands are essential for fine-grained manipulation, but their hardware designs vary substantially across embodiments.
arXiv:2606. 12109v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable zero-shot generalization in robotic manipulation, yet the vast majority of pre-trained pipelines remain strictly confined to low-DoF parallel grippers.
arXiv:2608. 15917v1 Announce Type: cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers.
arXiv:2606. 31694v1 Announce Type: cross Abstract: For robots manipulating open-world objects, tactile representations must generalize to unseen materials.
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.