arXiv Machine Learning

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning

arXiv:2511. 14427v4 Announce Type: replace-cross Abstract: Effective contact-rich manipulation requires robots to synergistically leverage vision, force, and proprioception.

arXiv AI
Aug 26

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

ForceFlow is a force-aware reactive framework that uses flow matching to improve contact-rich manipulation. It fuses force signals asymmetrically, treats force as a global regulator, and employs a joint prediction paradigm to couple force and motion. The approach splits tasks into a vision-dominant localization stage and a touch-dominant execution stage, using a Vision-to-Force handover to separate spatial generalization from contact regulation.

By Shuoheng Zhang, Yifu Yuan, Hongyao Tang, Yan Zheng, Qiaojun Yu, Pengyi Li, Guowei Huang, Helong Huang, Xingyue Quan, Jianye Hao
arXiv Machine Learning
Sep 10

Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers

The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.

By Sarvesh Patil
arXiv AI
Aug 20

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

ADEPT is a reinforcement‑learning framework that first pre‑trains a dexterous policy on a generic object reposing task and then post‑trains downstream policies using this pretrained behavior as a prior. The approach avoids relearning basic skills for each new task, and employs a stable post‑training recipe—behavior‑cloning distillation, critic warm‑up, and conservative on‑policy updates—to preserve the pretrained capabilities. ADEPT’s joint‑space Geometric Fabric mediates between the policy and the robot, enabling zero‑shot sim‑to‑real transfer on a 23‑DoF Kuka‑Allegro and a 29‑DoF Flexiv‑Sharpa, where the robots solve long‑horizon tasks from challenging initial states at human‑level speed.

By Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa
arXiv Computer Vision
Sep 21

ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.

By Xialin He, Sirui Xu, Xinyao Li, Runpei Dong, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
arXiv AI
Sep 23

Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors

Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.

By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
arXiv Machine Learning
Aug 4

DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration

arXiv:2608. 01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that are moving or require rapid adjustments.

By Haoran Liao, Pengyue Wang, Shuoyu Chen, Kehan Cheng, Xuhang Chen, Yuhao Lin, Mu Lin, Zhizhao Liang, Xiaoyi Fan, Chengyi Xing, Dan Niu, Yi-Lin Wei, Wei-Shi Zheng