Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv Computer Vision
23h ago

Native Action-Prior Learning from Videos for World Action Models

arXiv:2610.03391v1 Announce Type: new Abstract: World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotat...

By Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost, Zijian Zhou, Yikai Wang, Xudong Wang, Aditya Patel, Belinda Zeng, Tao Xiang, Serge Belongie, Amir Bar, Sen He
arXiv Computer Vision
23h ago

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

arXiv:2610.02521v1 Announce Type: new Abstract: Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observati...

By Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang
arXiv Machine Learning
23h ago

Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies

The paper examines how transformer policies, which process agents as ordered token sequences, perform in multi‑agent robot learning where agent teams are unordered. It finds that low permutation error can mask action collapse, where all agents choose the same action, and proposes additional diagnostics such as action diversity and same‑action fraction. Experiments show that while a PPO‑ID baseline avoids collapse, it remains order‑sensitive, and that strong equivariance regularization can still cause homogeneous behavior; a weak penalty improves robustness and preserves diversity for three‑agent teams, but four‑agent teams need much smaller regularization weights.

By Amit Thakur, Mukesh Singhal
arXiv Computer Vision
23h ago

EVEWorld: Physical Evolution Supervision for Embodied World Models

EVEWorld introduces a physical evolution-supervision framework for embodied world models, addressing the issue of Model Laziness by focusing on physical consistency rather than visual fidelity. The framework comprises Instance-Guided Restoration (IGR) to enforce instance consistency and Temporal Instance Alignment (TIA) to align target instances across adjacent frames. Experiments on DreamGenBench, EWMBench, and PBench show an 87.5% reduction in the Model Laziness Rate (MLR) compared to GigaWorld-0, and the model ranks 6th in JEPA Similarity on the WorldArena 2.0 Track 1 leaderboard.

By Kaiqi Wang, Songxin Zhang, Zejian Xie, Xiao Xiong, Zhuoyang Song, Ziwei Wu, Jun Yu Lu, Yitan Teng, Ziying Song, Jiaxing Zhang
arXiv Computation and Language
23h ago

How Robust Is Multimodal Claim Verification to LLM Rewriting?

The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.

By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
arXiv AI
23h ago

How Causality Bridges the Semantic Gap

The paper introduces CausalBridge, a framework that uses causal structure to assign semantics to unnamed variables in numerical measurements. By discovering a causal graph from the data and solving for variable embeddings constrained by this graph, the method aligns variable meanings with a language model, outperforming association‑based approaches. Experiments on questionnaires and robotics scenarios show high accuracy even when most variable names are masked, enabling rapid and cost‑effective system naming.

By Shuhao Zhang, Xuran Zhou, Han Guo, Pengtao Xie, Yujia Zheng
arXiv AI
23h ago

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

FastOPD is a framework that distills large Vision‑Language‑Action (VLA) models into lightweight versions by using on‑policy distillation with a flow map and a self‑consistency objective. The method trains a compact student to mimic the teacher’s dynamics, achieving performance close to the teacher while drastically reducing inference steps. Experiments on LIBERO, RoboTwin 2.0, and real‑robot deployments show significant latency reductions and higher success rates compared to existing few‑step distillation baselines.

By Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
arXiv Machine Learning
23h ago

How Far Does a Shared Linear Map Go? Probing Feature-Space Manipulability for Image Editing

The paper investigates how well a simple spatially shared linear map can predict changes in a model’s internal feature space caused by various image manipulations, including geometric, photometric, occlusion, and diffusion-generated semantic edits. Experiments across ConvNeXt, SwinV2, and DINOv3 show that this linear operator often performs nearly as well as more complex, higher‑capacity probes, especially in deeper layers of supervised backbones. The study concludes that a shared linear map is frequently sufficient to capture diverse image edits, with its leading singular components encoding semantic content and higher‑rank components refining details.

By Elias Krey, Nils Neukirch, Nils Strodthoff