arXiv AI By Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong We, Qiaojun Yu, Dibo Hou

FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

Read the original on arXiv AI →

FWBC‑VLA is a force‑aware framework that links vision‑language‑action (VLA) models with whole‑body compensation control for wheeled‑legged robots. It introduces HSR‑Force, a sensorless residual‑torque estimator that infers contact strength and encodes this information as tokens for the VLA action decoder, allowing the policy to detect contact onset, loading, and release. The system fine‑tunes a pretrained VLA backbone on a large WL&Arm dataset, combines proprioceptive, Jacobian‑derived force, and contact estimates to generate corrective actions, and demonstrates effectiveness in real‑world tasks such as whiteboard wiping and door opening.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

ForceFlow is a force-aware reactive framework that uses flow matching to improve contact-rich manipulation. It fuses force signals asymmetrically, treats force as a global regulator, and employs a joint prediction paradigm to couple force and motion. The approach splits tasks into a vision-dominant localization stage and a touch-dominant execution stage, using a Vision-to-Force handover to separate spatial generalization from contact regulation.

By Shuoheng Zhang, Yifu Yuan, Hongyao Tang, Yan Zheng, Qiaojun Yu, Pengyi Li, Guowei Huang, Helong Huang, Xingyue Quan, Jianye Hao
Hugging Face Trending Papers
Jun 3

VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions.