arXiv AI

X-Planner: Event-Structured Task Planning for Embodied Intelligence

arXiv AI
Jun 26

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

arXiv:2606. 27251v1 Announce Type: cross Abstract: Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation.

By Junhao Shi, Zezheng Huai, Siyin Wang, Jia Chen, Yubang Wang, Zhaoye Fei, Hechang Chen, Jingjing Gong, Xipeng Qiu, Yu-Gang Jiang
arXiv AI
Jul 7

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

arXiv:2607. 05377v1 Announce Type: cross Abstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations.

By Jiaqi Peng, Xiqian Yu, Delin Feng, Yuqiang Yang, Wenzhe Cai, Jing Xiong, Ganlin Yang, Jinliang Zheng, Jiafei Cao, Xueyuan Wei, Jiangmiao Pang, Yuan Shen, Tai Wang
Hugging Face Trending Papers
Jun 25

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (APIs, IoT) and physical (manipulation, navigation) domains, coupled with autonomous recovery from physical failures that inevitably arise over extended operation. Existing systems treat these as separate problems: VLM-based planners lack a unified cyber-physical action space, agent frameworks accumulate unbounded context that degrades temporal coherence, and VLA policies execute open-loop without detecting their own failures.

arXiv AI
Jul 28

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.

By Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo
arXiv AI
Jul 17

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

arXiv:2607. 14187v1 Announce Type: new Abstract: Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved.

By Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang
arXiv AI
Aug 28

STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration

The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.

By Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
arXiv Computer Vision
2d ago

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.

By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang