ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 16202v1 Announce Type: cross Abstract: Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformable dynamics, such as elastic materials and fabrics, remains a major challenge for computer vision and robotics.
FWBC‑VLA is a force‑aware framework that links vision‑language‑action (VLA) models with whole‑body compensation control for wheeled‑legged robots. It introduces HSR‑Force, a sensorless residual‑torque estimator that infers contact strength and encodes this information as tokens for the VLA action decoder, allowing the policy to detect contact onset, loading, and release. The system fine‑tunes a pretrained VLA backbone on a large WL&Arm dataset, combines proprioceptive, Jacobian‑derived force, and contact estimates to generate corrective actions, and demonstrates effectiveness in real‑world tasks such as whiteboard wiping and door opening.
arXiv:2609.37067v1 Announce Type: cross Abstract: Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Constr...
arXiv:2607. 19190v1 Announce Type: cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation.
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across visual perception tools and simulators.
arXiv:2609.15726v1 Announce Type: cross Abstract: Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not conve...