arXiv AI

MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

MorphIK is a flow‑matching neural model that learns inverse kinematics for revolute‑joint kinematic chains it has never seen during training. Using a transformer to encode a robot’s morphology and target pose, the model generates pose solutions from noise and can be fine‑tuned with optimization to achieve sub‑centimeter accuracy. It also efficiently samples the robot’s null space, producing diverse configurations for the same pose.

arXiv Machine Learning
Jun 9

RAM: Reachability Across Morphologies

arXiv:2606. 09108v1 Announce Type: cross Abstract: Many stages of the robotic lifecycle, from morphology synthesis to operation, rely fundamentally on the reachable workspace.

By Tim Walter, Xinyu Chen, Jonathan K\"ulz, Matthias Althoff
arXiv Machine Learning
Jun 17

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

arXiv:2606. 17846v1 Announce Type: cross Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale.

By Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen
arXiv Machine Learning
Aug 31

FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation

FlowCorrect is a modular interactive imitation learning method that allows real‑time adaptation of generative flow‑matching manipulation policies using sparse, relative human corrections. During task execution, a human provides brief corrective pose nudges through a lightweight VR interface, and FlowCorrect locally adapts the policy without retraining the backbone, maintaining performance on previously learned scenarios. Experiments on a real‑world robot across four tabletop tasks show that, with a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on solved scenarios.

By Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
arXiv Computer Vision
Sep 25

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.

By Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao
arXiv AI
Jul 1

A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

arXiv:2509. 15443v2 Announce Type: replace-cross Abstract: Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing abundant, large-scale human motion collections.

By Xingyu Chen, Hanyu Wu, Sikai Wu, Mingliang Zhou, Diyun Xiang, Haodong Zhang, Yangchen Zhou, Yukang Gao, Yi Gu, Renjing Xu
arXiv AI
Sep 24

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.

By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
arXiv AI
Sep 4

BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI

The paper introduces BRIDGE, an open‑source 88 cm tall humanoid robot designed through a data‑driven morphology‑control co‑design framework that optimizes the robot’s body shape for human‑like movement. A new metric combining kinematic retargeting fidelity and dynamic tracking performance is proposed to evaluate morphological fidelity, and the framework achieves state‑of‑the‑art results compared to existing humanoids such as Bumi, K1, and Toddlerbot. The resulting platform, released with its control policy and supporting materials, demonstrates superior fidelity in capturing human motion, robust balance, and highly dynamic maneuvers.

By Jianren Wang, Letian Qian, Zikai Wang, Weiwei Wu, Junjie Zong, Abhinav Gupta, Deepak Pathak
arXiv Computer Vision
Sep 15

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv:2604.28130v4 Announce Type: replace Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts join...

By Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu, Weixia He, Qi Wang, Ning Zhang, Zhengyu Li, Guanli Hou, Dongze Lian, Xiaoyu He, Mingyuan Zhang, Hanwang Zhang