ForceFlow is a force-aware reactive framework that uses flow matching to improve contact-rich manipulation. It fuses force signals asymmetrically, treats force as a global regulator, and employs a joint prediction paradigm to couple force and motion. The approach splits tasks into a vision-dominant localization stage and a touch-dominant execution stage, using a Vision-to-Force handover to separate spatial generalization from contact regulation.
By Shuoheng Zhang, Yifu Yuan, Hongyao Tang, Yan Zheng, Qiaojun Yu, Pengyi Li, Guowei Huang, Helong Huang, Xingyue Quan, Jianye Hao
arXiv:2608. 10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency and generalization in robot learning.
By Jusuk Lee, Daesol Cho, Jonghun Shin, Seungyeon Yoo, Jonghae Park, Taekbeom Lee, H. Jin Kim
Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging.
The paper introduces Distributed Dexterous Manipulation (DDM), a challenging control problem involving 64 soft delta robots arranged in an 8x8 grid. It presents a framework using spatially conditioned Multi-Agent Transformers (MATs) with adaptive layer norm, spatial contrastive embeddings, and a behavior cloning method fine‑tuned by Soft Actor Critic. Experiments demonstrate that MATs refine actions through stacked attention blocks, enabling long‑horizon planar manipulation in simulation and real‑world settings, while an action‑selection strategy reduces robot usage by about 65% and lowers wear‑and‑tear, achieving an average error of ~1.5 cm.
By Sarvesh Patil
arXiv:2604. 20348v2 Announce Type: replace-cross Abstract: Language Models (LLMs) have emerged as powerful reasoning engines for embodied control.
By Alessio Palma, Indro Spinelli, Vignesh Prasad, Luca Scofano, Yufeng Jin, Georgia Chalvatzaki, Fabio Galasso
arXiv:2607. 27549v1 Announce Type: cross Abstract: Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments.
By Ajay Sridhar, Jensen Gao, Jonathan Yang, Jean Mercat, Suneel Belkhale, Dorsa Sadigh
ADEPT is a reinforcement‑learning framework that first pre‑trains a dexterous policy on a generic object reposing task and then post‑trains downstream policies using this pretrained behavior as a prior. The approach avoids relearning basic skills for each new task, and employs a stable post‑training recipe—behavior‑cloning distillation, critic warm‑up, and conservative on‑policy updates—to preserve the pretrained capabilities. ADEPT’s joint‑space Geometric Fabric mediates between the policy and the robot, enabling zero‑shot sim‑to‑real transfer on a 23‑DoF Kuka‑Allegro and a 29‑DoF Flexiv‑Sharpa, where the robots solve long‑horizon tasks from challenging initial states at human‑level speed.
By Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa
arXiv:2607. 15880v1 Announce Type: cross Abstract: Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization.
By Zhenduo Shang, Xiyao Liu, Bohan Li, Xudong Wang, Teng Ren, Lianqing Liu, Zhi Han
arXiv:2606. 06218v1 Announce Type: cross Abstract: A policy tuned for one robot often behaves differently on another, whether due to the sim-to-real gap, unknown payloads, or the differing dynamics of two instances of the same robot.
By Dongwon Son, Florian Shkurti, Jason Lee, Naman Shah, Beomjoon Kim, Dieter Fox
ULTRA is a unified framework for autonomous humanoid whole-body locomotion and manipulation that overcomes limitations of prior methods by combining a physics-driven neural retargeting algorithm with a multimodal controller. The retargeting algorithm translates large-scale motion capture data into physically plausible humanoid motions, while the controller learns to handle both dense motion references and sparse task specifications using a range of sensory inputs, from accurate motion-capture states to noisy egocentric vision. In simulation and on a real Unitree G1 humanoid, ULTRA demonstrates improved generalization and robustness, enabling coordinated whole-body behavior from sparse intent without relying on test-time reference motions.
By Xialin He, Sirui Xu, Xinyao Li, Runpei Dong, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
Tactile-JEPA is a self‑supervised pre‑training method for distributed tactile sensors that leverages the sensors’ spatial topology to learn topology‑aware representations. It predicts embeddings of masked sensing elements using a sensor connectivity graph and dual‑scale masking to capture both local contact details and the global tactile surface state. Evaluated on three diverse datasets, it improves force estimation by 6.3 % and in‑hand orientation error by 20.8 % over previous state‑of‑the‑art methods, and yields consistent gains in downstream tasks such as policy learning.
By Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy
arXiv:2608. 01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that are moving or require rapid adjustments.
By Haoran Liao, Pengyue Wang, Shuoyu Chen, Kehan Cheng, Xuhang Chen, Yuhao Lin, Mu Lin, Zhizhao Liang, Xiaoyi Fan, Chengyi Xing, Dan Niu, Yi-Lin Wei, Wei-Shi Zheng