arXiv:2610.03537v1 Announce Type: cross
Abstract: Endovascular brain-computer interfaces (BCIs) avoid craniotomy but require precise device delivery through anatomically variable cerebral veins. This...
By Harry Robertshaw, Weijie Qi, Nikola Fischer, Alejandro Granados, Thomas C. Booth, Sam E. John
arXiv:2610.03273v1 Announce Type: new
Abstract: Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concep...
By Geonwoo Bang, Dongho Kim, Moohong Min
arXiv:2610.02974v1 Announce Type: cross
Abstract: Image-based traversability estimation is inherently dependent on the robot platform, deployment domain, and mission preferences, which limits the app...
By Simon Schwaiger, David Seyser, Alessandro Scherl, Zlatan Ajanovi\'c, Wilfried W\"ober, Gerald Steinbauer-Wagner
arXiv:2610.03391v1 Announce Type: new
Abstract: World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotat...
By Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost, Zijian Zhou, Yikai Wang, Xudong Wang, Aditya Patel, Belinda Zeng, Tao Xiang, Serge Belongie, Amir Bar, Sen He
arXiv:2610.03015v1 Announce Type: cross
Abstract: Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors....
By Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang
arXiv:2610.02662v1 Announce Type: new
Abstract: Language-enabled robot systems increasingly combine semantic-graph planning with temporal-logic safety monitors. We investigate a trace-completeness as...
By Stabak Das, Priyesh Ranjan, Xiangfang Li, Lijun Qian
arXiv:2610.02521v1 Announce Type: new
Abstract: Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observati...
By Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang
arXiv:2610.02825v1 Announce Type: new
Abstract: Road geometry (e.g., crests, sags, and speed humps) and surface conditions (e.g., wet or icy pavement) affect how vehicles move, what drivers and onboa...
By Yang Chen, Yicheng Zhu, zhenning Li, Tao Li, Zilin Bian
arXiv:2607.27599v2 Announce Type: replace
Abstract: Building generalizable robot agents for diverse applications remains a fundamental challenge. While imitation learning-based policies can perform w...
By Xiangcheng Zhang, Runhan Huang, Yilun Du
arXiv:2610.03664v1 Announce Type: new
Abstract: Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive par...
By Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen
arXiv:2512.01352v2 Announce Type: replace
Abstract: Unsupervised and open-vocabulary 3D object detection have recently gained attention, particularly in autonomous driving, where reducing annotation...
By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
arXiv:2610.02242v1 Announce Type: cross
Abstract: Chemical reaction-condition optimization -- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize...
By Lingli Ge, Yubin Wang, Junyuan Gao, Jiahe Song, Jiaxing Sun, Boyu Zhu, Haote Yang, Jingchao Wang, Lixin Ma, Jiang Wu, Yuqiang Li, Conghui He
arXiv:2610.02542v1 Announce Type: new
Abstract: World models (WMs) simulate the transition dynamics of environments, enabling agents to plan over the consequences of their actions. In text-based envi...
By Dhananjay Ashok, Shantanu Agarwal, Vivek Datla, Jonathan May, Alfy Samuel
arXiv:2610.02366v1 Announce Type: new
Abstract: Finding an optimal behaviour policy within a given environment is a widely studied problem in domains as diverse as games, robotics, energy infrastruct...
By Aviraj Newatia, Yordan Tsvetkov, Leonard Pleiss, Andrew Spielberg, Rika Antonova
The paper examines how transformer policies, which process agents as ordered token sequences, perform in multi‑agent robot learning where agent teams are unordered. It finds that low permutation error can mask action collapse, where all agents choose the same action, and proposes additional diagnostics such as action diversity and same‑action fraction. Experiments show that while a PPO‑ID baseline avoids collapse, it remains order‑sensitive, and that strong equivariance regularization can still cause homogeneous behavior; a weak penalty improves robustness and preserves diversity for three‑agent teams, but four‑agent teams need much smaller regularization weights.
By Amit Thakur, Mukesh Singhal
EVEWorld introduces a physical evolution-supervision framework for embodied world models, addressing the issue of Model Laziness by focusing on physical consistency rather than visual fidelity. The framework comprises Instance-Guided Restoration (IGR) to enforce instance consistency and Temporal Instance Alignment (TIA) to align target instances across adjacent frames. Experiments on DreamGenBench, EWMBench, and PBench show an 87.5% reduction in the Model Laziness Rate (MLR) compared to GigaWorld-0, and the model ranks 6th in JEPA Similarity on the WorldArena 2.0 Track 1 leaderboard.
By Kaiqi Wang, Songxin Zhang, Zejian Xie, Xiao Xiong, Zhuoyang Song, Ziwei Wu, Jun Yu Lu, Yitan Teng, Ziying Song, Jiaxing Zhang
The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.
By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
The paper introduces CausalBridge, a framework that uses causal structure to assign semantics to unnamed variables in numerical measurements. By discovering a causal graph from the data and solving for variable embeddings constrained by this graph, the method aligns variable meanings with a language model, outperforming association‑based approaches. Experiments on questionnaires and robotics scenarios show high accuracy even when most variable names are masked, enabling rapid and cost‑effective system naming.
By Shuhao Zhang, Xuran Zhou, Han Guo, Pengtao Xie, Yujia Zheng
FastOPD is a framework that distills large Vision‑Language‑Action (VLA) models into lightweight versions by using on‑policy distillation with a flow map and a self‑consistency objective. The method trains a compact student to mimic the teacher’s dynamics, achieving performance close to the teacher while drastically reducing inference steps. Experiments on LIBERO, RoboTwin 2.0, and real‑robot deployments show significant latency reductions and higher success rates compared to existing few‑step distillation baselines.
By Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park, Seonghyun Jin, Sunwoo Park, Youngmin Kim, Youngjun Jun, Kyumin Choi, Jong Chul Ye
The paper investigates how well a simple spatially shared linear map can predict changes in a model’s internal feature space caused by various image manipulations, including geometric, photometric, occlusion, and diffusion-generated semantic edits. Experiments across ConvNeXt, SwinV2, and DINOv3 show that this linear operator often performs nearly as well as more complex, higher‑capacity probes, especially in deeper layers of supervised backbones. The study concludes that a shared linear map is frequently sufficient to capture diverse image edits, with its leading singular components encoding semantic content and higher‑rank components refining details.
By Elias Krey, Nils Neukirch, Nils Strodthoff