Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv Computer Vision
4d ago

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

arXiv:2610.01741v1 Announce Type: new Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, exist...

By Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
arXiv Computer Vision
4d ago

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

arXiv:2610.01939v1 Announce Type: new Abstract: Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant obser...

By Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou
arXiv Computer Vision
4d ago

A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform

The paper reviews end‑to‑end autonomous driving (E2E‑AD) training, framing it as a Data‑Strategy‑Platform system. It surveys recent advances in data pipelines, learning paradigms, and training infrastructures, and discusses how these layers interact to influence model performance, robustness, and deployability. The authors highlight current limitations and propose a future vision that prioritizes data value, foundation‑driven generalization, and integrated training‑testing loops for more robust, scalable, and trustworthy autonomous driving systems.

By Chengkai Xu, Yiming Cui, Jiaqi Liu, Yicheng Guo, Cheng Qin, Geyuan Zhang, Xinwei Dong, Shiyu Fang, Peng Hang, Jian Sun
arXiv Computer Vision
4d ago

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.

By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv AI
4d ago

Network World Models as Environments for Algorithm Design on Complex Systems

The paper introduces an action‑conditioned Network World Model that learns how a network’s diffusion dynamics evolve under interventions over time. This model can quickly predict the outcomes of actions, enabling a coding agent to design and refine algorithms that select actions to maximize expected performance on complex network tasks. Experiments on eight network tasks and five diffusion models show that the resulting algorithms match or surpass the best existing baselines in 138 of 141 settings while achieving up to 14.5× faster rollouts than traditional Monte Carlo simulation.

By Rishab Alagharu, Hongji Pu, Zeeshan Memon, Xinyuan Song, Yuntong Hu, Liang Zhao
arXiv AI
4d ago

Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks

The paper introduces a bounded‑fidelity sim‑as‑demo‑stage design pattern that suppresses contact physics during object handoffs in simulators, using MuJoCo’s mocap‑body primitive and a lightweight Python adapter. This approach ensures audit‑chain stability, producing identical event‑log hashes across 1,000 replays per posture, whereas a contact‑force baseline yields significant divergence. The authors demonstrate that the pattern maintains reproducibility across various timesteps and sequential handoffs with minimal overhead, and they identify contexts where it should not be applied.

By Xue Qin, Simin Luan, Cong Yang, Zhijun Li
arXiv AI
4d ago

Probabilistic Plan Legibility with Off-the-shelf Planners

The paper introduces a method for generating legible plans in arbitrary PDDL domains by extending prior legibility research to classical planning without custom planners. It incorporates a second‑order theory of mind to estimate the observer’s perspective, enabling robots to implicitly communicate goals in human‑robot teaming. Benchmark results show that increasing legibility typically trades off with plan efficiency, and a regularizing factor is needed to balance the two.

By Michele Persiani, Thomas Hellstr\"om