Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
5d ago

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.

By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
arXiv AI
5d ago

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.

By Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao
arXiv AI
5d ago

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

HiWE is a hierarchical world knowledge model that enables zero‑shot 3D path planning by linking visual grounding with language‑based planning through a point‑based interface. It uses PointVLM to map task‑relevant objects to image coordinates, lifts these predictions into a semantic 3D representation with depth data, and then a language planner (3DLLM) generates end‑effector waypoints and gripper commands. The system is evaluated on 14 simulated manipulation tasks and four physical‑robot tasks, with ablations on visual training data, spatial inputs, and grasp selection.

By Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu
arXiv AI
5d ago

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

GroundingPI is a 4‑billion‑parameter grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It is trained with multimodal and spatial pretraining, supervised fine‑tuning, and reinforcement learning, achieving a new state‑of‑the‑art average of 73.68% across 34 grounding benchmarks. As a visual backbone, GroundingPI improves performance in robotic manipulation and autonomous driving, outperforming larger models and mainstream backbones in several out‑of‑distribution settings.

By Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
arXiv AI
5d ago

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

RoboCoach introduces a world-model-guided coaching framework that uses imagined failures to direct demonstration requests and expert updates for robot manipulation tasks. The system, called RIDI, runs reusable skill experts within a shared action-conditioned world model and records the first failing subtask to decide which demonstrations to acquire and which adapters to refine. Experiments on simulation suites and real robots show that with only 150 additional subtask demonstrations, success rates increase dramatically, and the coached experts transfer effectively to unseen task compositions.

By Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
arXiv AI
5d ago

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.

By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
arXiv AI
5d ago

BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph

BatSLAM 2.0 is a sonar-only SLAM system that builds a topological map by recognizing places from acoustic signals. It incorporates an updated acoustic front-end, a sequence verifier that tracks and verifies loop closure candidates, and a pose graph implemented on a high-performance factor graph framework. The system was thoroughly evaluated in both simulated and real-world recordings, demonstrating robust topological map creation, prevention of map collapse, and scalable map size.

By Jan Steckel
arXiv AI
5d ago

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

PrefPI (Preference-Guided Policy Iteration) is an iterative framework that steers pretrained generative robot policies using only relative preferences over self-generated trajectories. It treats preference learning as preference-conditioned generative modeling, where preferred trajectories define a conditional distribution whose density ratio with the broader behavior prior yields an implicit preference signal amplified by classifier-free guidance (CFG). By repeatedly applying this preference-conditioned modeling and guidance, PrefPI iteratively improves policies, enabling access to behaviors that were rarely or never observed under the initial policy, and achieves significant behavioral shifts such as increasing object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.

By Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha
arXiv AI
5d ago

X-Planner: Event-Structured Task Planning for Embodied Intelligence

arXiv:2609.25187v2 Announce Type: replace Abstract: Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems...

By Howard Lu, Shalfun Li, Porter Pan, Cris, Lumen, Cyril, Eric Hu, Lily Li, Maeve Zhang, Rain Sun, Robert Wang, KZ Zheng, Viggo Chen, Tim Ding, Regsis Cheng, YJ Xiao, Kian, Hai Lin, Alan Song, Elise Ma, Gody Li, Victor Yao, Yohann Tang, Ingrid Yu, Jason He, James Wang, Ryan Yu, Ping Yang, Chris Pan, Vincent Chen, Roy Gan, Hao Wang, Qian Wang
arXiv Computation and Language
5d ago

Spike-driven Vision-Language-Action Model

arXiv:2609.39514v1 Announce Type: new Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...

By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
arXiv Computation and Language
5d ago

TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns

arXiv:2609.38371v1 Announce Type: cross Abstract: Existing LLM-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focu...

By Guangxin Zhao, Yiran Hu, Yuan Cao, Chenxi Jiang, Jianfei Yang, Yegang Du, Yasuyuki Taki, Yoshifumi Kitamura, Lin Gu, Zhi Zheng
arXiv Computer Vision
5d ago

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

arXiv:2609.38641v1 Announce Type: new Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...

By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone