The Potential of Haptic Foundation Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.29601v1 Announce Type: cross Abstract: We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale m...
arXiv:2608.29601v2 Announce Type: replace-cross Abstract: We present $N_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale mul...
arXiv:2603. 14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models.
arXiv:2606. 31451v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities.
arXiv:2606. 11637v1 Announce Type: new Abstract: Touch is a key modality for embodied agents to understand the physical world.
VT-MUSE is a multimodal unified sequential representation learning framework for visuotactile manipulation. It uses a two‑stage approach: first, modality‑specific encoders are jointly adapted with cross‑modal temporal alignment and masked‑view consistency; second, a conditional variational latent model processes masked visual sequences and full tactile histories, with auxiliary decoders reconstructing recent visual observations and predicting tactile depth changes. The resulting representation is fed into a lightweight Transformer policy via gated cross‑attention, achieving an 11‑percentage‑point improvement over the strongest baseline in simulation and significant gains in real‑world experiments.