Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

4,108 stories · RSS feed

arXiv AI
Sep 2

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

EEG-VID is a task‑guided latent predictive pretraining framework designed to improve EEG decoding across session and subject shifts. It predicts future latent EEG states from recent history using an exponential‑moving‑average target encoder and weak task guidance, then fine‑tunes with supervised learning. The method achieves significant accuracy gains on VIG‑48 and BCI Competition datasets, and demonstrates effective assistive target selection in a robot‑scene study.

By Guanzhong Sun, Junyi Ma, Yuxuan Wu, Yanzi Miao
arXiv Computer Vision
Sep 2

Hydra: Marker-Free RGB-D Hand-Eye Calibration

Hydra introduces a marker‑free RGB‑D hand‑eye calibration method that leverages a novel ICP algorithm with a robust point‑to‑plane objective on a Lie algebra. Experiments on three serial manipulators and two RGB‑D cameras show that with only three random robot configurations the method achieves about 90% successful calibrations, 2–3× faster convergence to the global optimum, and 2 orders of magnitude faster convergence time (0.8 ± 0.4 s) compared to other marker‑free baselines. The approach delivers improved accuracy (5 mm in task space versus 7 mm for classical methods) while remaining marker‑free, and the authors provide an open‑source dataset, code, and ROS 2 integration.

By Martin Huber, Huanyu Tian, Christopher E. Mower, Lucas-Raphael M\"uller, S\'ebastien Ourselin, Christos Bergeles, Tom Vercauteren
arXiv Computation and Language
Sep 2

Apples on the Table? Evaluating Text-Guided 3D Scene Synthesis via Fine-Grained Constraint Verification

The paper introduces LEGO, a benchmark dataset pairing user text descriptions with human‑annotated fine‑grained constraints and reference 3D scenes, and LEGO‑Eval, an evaluation framework that decomposes descriptions into atomic constraints and verifies each using grounding and spatial reasoning tools. It demonstrates that LEGO‑Eval detects misalignment more accurately than existing methods and that current 3D scene synthesis approaches achieve at most a 10% success rate on this benchmark.

By Minseok Kang, Dongwook Choi, Gyeom Hwangbo, Seungwon Lim, Kai Tzu-iunn Ong, Jinyoung Yeo
arXiv AI
Sep 2

Dual Process Motion Planning

The paper introduces a dual‑process architecture for nonlinear motion planning that blends fast, learning‑based intuition (System‑1) with slow, robust symbolic reasoning (System‑2). A metacognitive controller decides when to use each component, aiming to balance speed, precision, and adaptability. Experiments on diverse benchmark environments show consistent improvements in planning efficiency, accuracy, and generalization, highlighting the benefits of integrating learning with structured reasoning.

By Jiayi Yan, Francesco Fabiano, Alessandro Abate
arXiv AI
Sep 2

Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting

The paper introduces a cross‑modal pseudo‑labeling pipeline for unsupervised domain adaptation in semantic segmentation, particularly for waste sorting. It combines SAM for class‑agnostic region proposals with EVA‑CLIP to assign semantic labels via region‑text similarity, applying confidence filtering to ensure reliable pseudo‑labels for self‑training. An optional BLIP‑based language‑grounded verification further refines ambiguous regions, and the method shows consistent improvements over source‑only baselines on synthetic‑to‑real driving and lab‑to‑factory waste sorting shifts.

By Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl
arXiv AI
Sep 2

CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction

The paper introduces CoLT-Drive, a 3,536-sample counterfactual long‑tail benchmark for evaluating decision‑level driving affordance prediction, which tests whether models can infer how rare objects affect an ego vehicle’s high‑level actions. It also proposes KPA, a knowledge‑preserving adaptation framework that combines structured prompting, expert merging, and a regime‑aware LoRA mixture‑of‑experts module to improve small VLMs on driving tasks. Experiments show KPA achieves 60.8% pair accuracy on CoLT‑Drive, outperforming the Qwen3‑VL‑2B baseline and LoRA SFT while keeping competitive in‑domain performance.

By Zhengxu Tang, Guofeng Cui, Ziyu Gong, Xiaozhou Zhang, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang
arXiv AI
Sep 2

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.

By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
arXiv Computer Vision
Sep 2

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Qwen-Drive-1.0 is a vision‑language foundation model tailored for autonomous driving that builds on a pretrained VLM architecture. It incorporates a bird’s‑eye‑view perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert that generates future ego trajectories from shared representations. Experiments show strong 3D perception, driving scene understanding, and competitive motion‑planning performance while largely preserving general vision‑language capabilities.

By Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
arXiv Computer Vision
Sep 2

Online camera-pose-free stereo endoscopic tissue deformation recovery with tissue-invariant vision-biomechanics consistency

The paper presents a camera‑pose‑free stereo endoscopic method for recovering tissue deformation by modeling geometry as a 3D point‑derivative map and deformation as a 3D displacement‑local deformation map. It optimizes inter‑frame deformation in a camera‑centric setting, eliminating the need for camera pose estimation, and introduces a canonical map for online geometry and deformation optimization. Experiments on in‑vivo and ex‑vivo laparoscopic data show accurate 3D reconstruction (≈0.37–0.39 mm surface distance) even under occlusion, and the method can estimate surface strain distributions during manipulation.

By Jiahe Chen, Naoki Tomii, Ichiro Sakuma, Etsuko Kobayashi
arXiv AI
Sep 2

When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency

The paper examines how responsibility is assigned when AI systems fail, proposing a sociotechnical theory that distinguishes between AI incidents, organisational crises, and scandals. It argues that the configuration of an incident shapes actor-specific attribution, which in turn influences perceptions of capability, integrity, fairness, and relationships, and that public moralisation can elevate an incident to scandal. The authors introduce ‘accountable transparency’—a response framework combining timely notice, intelligible accounts, role acknowledgement, remedy, evidence of correction, and recourse—as a way to manage blame, trust, and communication credibility.

By Mohammad Saleh Torkestani, Taha Mansouri
arXiv AI
Sep 2

Towards Generalizable Visually Grounded Exploration of Household Devices

The paper introduces VGEBench, a new benchmark for evaluating Vision‑Language Models (VLMs) on generalizable, visually grounded exploration of household devices. Unlike existing datasets that rely on static images or annotated trajectories, VGEBench employs a logic‑driven state machine to simulate multi‑turn interaction loops, requiring agents to actively perceive, act, and refine their actions to achieve goals. Experiments show that current VLMs struggle to translate semantic knowledge into physical execution and to maintain long‑horizon state tracking.

By Linhao Zheng, Zeming Liu, Wangke Chen, Li Zeng, Wanxiang Che, Heyan Huang, Yuhang Guo
arXiv Computer Vision
Sep 2

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

CameraEditor is a new framework that transforms camera-controlled image editing into a temporal sequence prediction problem. By using video diffusion models, it incorporates a geometric perception module and dynamic reference routing to create precise visual references through dynamic panorama cropping. The method also inserts intermediate transition frames to handle large perspective shifts, maintaining content identity and spatial coherence, and is evaluated on a dataset of 5,760 instances with a benchmark of 462 test cases, achieving state‑of‑the‑art performance.

By Xin Shen, Chengyou Jia, Keshuo Xing, Zifeng Zhu, Changliang Xia, Bowen Ping, Zhuohang Dang, Hangwei Qian, Minnan Luo
arXiv AI
Sep 2

SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning

SEBA is a sample‑efficient framework for black‑box adversarial attacks on visual reinforcement learning agents. It combines a shadow Q model, a generative adversarial network for imperceptible perturbations, and a world model to simulate dynamics, reducing real‑world queries. Experiments on MuJoCo and Atari show SEBA significantly lowers cumulative rewards while preserving visual fidelity and requiring far fewer environment interactions than previous methods.

By Tairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye, Haibo Hu