arXiv AI

Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control

arXiv:2506. 08630v3 Announce Type: replace Abstract: A universal controller for any robot morphology would greatly improve computational and data efficiency.

arXiv AI
Jun 9

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

arXiv:2606. 08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments.

By Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan
arXiv AI
Jul 24

Emergent Compositional Skills in Mixture-of-Experts VLAs

arXiv:2607. 20771v1 Announce Type: cross Abstract: We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of task decomposition or hierarchy.

By Shlok Shah, Rhiaan Jhaveri, Tharun Kumar Tiruppali Kalidoss, Chirayu Nimonkar, Ishaan Javali
arXiv AI
Sep 16

QDTraj: Exploration of Diverse Trajectory Primitives for Articulated Objects Robotic Manipulation

The paper introduces QDTraj, a method that uses Quality‑Diversity algorithms to automatically generate a diverse set of low‑level trajectory primitives for manipulating articulated objects. By leveraging sparse reward exploration, QDTraj produces at least five times more diverse trajectories for hinge and slider tasks compared to baseline methods, and demonstrates strong generalization across 30 articulations from the PartNetMobility dataset, averaging 704 trajectories per task. The resulting primitives are validated both in simulation and on real robots, with the code released publicly.

By Mathilde Kappel, Mahdi Khoramshahi, Louis Annabi, Faiz Ben Amar, St\'ephane Doncieux
arXiv AI
Jul 28

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

arXiv:2607. 23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail.

By Daphne Chen, Archit Ritesh Jain, Eric Goossen, Emma Romig, Michael Murray, Nick Walker, Maya Cakmak
arXiv Computer Vision
3d ago

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv AI
Aug 19

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

LoopVLA introduces a recurrent Vision‑Language‑Action architecture that learns to refine multimodal representations, predict actions, and estimate when further refinement is unnecessary. By iteratively applying a shared Transformer block and producing a sufficiency score at each step, it decouples refinement from fixed layer indices and aligns confidence scores with action quality through a self‑supervised objective. Experiments on LIBERO, LIBERO‑Plus, and VLA‑Arena demonstrate that LoopVLA reduces model parameters by 45% and boosts inference throughput up to 1.7× while matching or surpassing strong baselines in task success.

By Boyang Shen, Kaixiang Yang, Hao Wang, Qiuyu Yu, Qiang Xie, Qiang Li, Zhiwei Wang
arXiv AI
Sep 25

MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

MorphIK is a flow‑matching neural model that learns inverse kinematics for revolute‑joint kinematic chains it has never seen during training. Using a transformer to encode a robot’s morphology and target pose, the model generates pose solutions from noise and can be fine‑tuned with optimization to achieve sub‑centimeter accuracy. It also efficiently samples the robot’s null space, producing diverse configurations for the same pose.

By Lennart Clasmeier, Jan Gerrit Habekost, Cornelius Weber, Stefan Wermter