arXiv:2609.38927v1 Announce Type: cross
Abstract: World models aim to learn representations of real-world environments and predict their future evolution. Recent object-centric world models have made...
By Yaqi Yang, Shuo Huang, Yujin Huang, Fucai Ke, Jiatong Han, Xin Zheng
arXiv:2609.40219v1 Announce Type: cross
Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experienc...
By Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
arXiv:2603.26741v2 Announce Type: replace-cross
Abstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, languag...
By Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Yetong Sha, Qiyu Hu, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng
arXiv:2604.11751v2 Announce Type: replace-cross
Abstract: World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the...
By Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li, Letian Wang, Alexandre Alahi, Harold Soh
arXiv:2605.15618v2 Announce Type: replace-cross
Abstract: Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reporte...
By Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar
arXiv:2609.39467v1 Announce Type: new
Abstract: Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-wo...
By ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu
arXiv:2609.40007v1 Announce Type: cross
Abstract: A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show...
By Yatharth Agarwal, Vijay Raghunathan
arXiv:2512.22819v2 Announce Type: replace
Abstract: Panoramic depth estimation captures the complete 360$^\circ$ scene geometry, being essential for robotics and AR/VR applications. While perspective...
By Hualie Jiang, Ziyang Song, Zhiqiang Lou, Rui Xu, Minglang Tan
arXiv:2609.34792v2 Announce Type: replace
Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
The paper introduces TacEx, a tactile‑curiosity framework that guides reinforcement learning agents to explore contact dynamics by focusing epistemic uncertainty on the tactile channel. By anchoring curiosity to touch, robots learn to manipulate and grasp objects without task rewards or demonstrations, generating an interaction‑dense dataset that supports offline pick‑and‑place policy learning. TacEx also enhances vision‑language‑action models through post‑training, improving downstream performance while remaining sample‑efficient.
By Klemens Iten, Alexander Proshkin, Bhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel, Carmelo Sferrazza
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a pr...
Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present...
Digital image watermarking is increasingly critical in media contexts, as emerging regulations and industry practices require marking AI-generated content and ensuring traceable sources to prevent man...
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly...
Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physica...
arXiv:2609.37239v1 Announce Type: new
Abstract: Many problems in sequential decision-making, such as imitation learning from observations, state-space compression, world-model learning, and sim-to-re...
By Sergio Calo, Amy Zhang, Javier Segovia-Aguas, Anders Jonsson
CAR‑VLA is a Vision‑Language‑Action model for autonomous driving that jointly considers scene complexity and dynamic risk to determine reasoning depth, urgency, and focus. It maps four complexity‑risk categories to three reasoning modes—Fast Intuition, Slow Thinking, and Reflex Response—each tailored to different driving scenarios. The model is trained via progressive supervised learning and reinforcement learning, achieving competitive performance on NAVSIM and Navhard benchmarks and demonstrating risk‑aware reasoning in high‑risk scenarios.
By Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.
By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
RawVLA introduces a streaming neural image signal processor that adaptively renders RAW observations for vision‑language‑action (VLA) policies, focusing on imaging factors that influence embodied behavior. The authors systematically analyze how five ISP dimensions—gain, sensor noise, chromatic response, tonal response, and bit depth—affect action prediction and manipulation success, showing that RAW‑to‑RGB processing significantly shapes outcomes. They also present RawVLA‑Bench, a RAW‑domain manipulation benchmark that evaluates image processing as an explicit variable across clean and adverse conditions, demonstrating that RawVLA maintains performance under standard settings while markedly improving robustness under degraded imaging.
By Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, Jianfei Yang, Ziteng Cui
CogWAM is a cognition‑guided world‑action model that links task reasoning with action learning via a persistent Semantic State, which records completed task events and the active subtask. The model updates this state only on semantic transitions, enabling task‑level context to persist across multiple action chunks. During inference, CogWAM removes the future‑prediction branch and directly generates actions from observations and the maintained state, achieving state‑of‑the‑art performance on RoboDojo and BiCoord and demonstrating efficient closed‑loop dual‑arm manipulation in real‑world experiments.
By Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen, Haoyi Jiang, Taojun Ding, Tingyang Xiao, Zhizhong Su, Jie Wang, Sanping Zhou