Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

4,108 stories · RSS feed

arXiv Computer Vision
Sep 7

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation

CrossDepth introduces geometry-constrained attention for multi-view surround depth estimation, addressing cross-image inconsistencies caused by varying camera intrinsics and limited receptive fields. The method conditions features on per-pixel camera-aware ray embeddings and extends pixel context via cross-image attention limited to geometrically plausible regions. Trained self-supervised with photometric consistency, it achieves better depth accuracy and consistency on DDAD and nuScenes compared to existing self-supervised approaches.

By Samer Abualhanud, Max Mehltretter
arXiv Computer Vision
Sep 7

WorldSculpt: Generating Compositional Worlds from Grounded Videos

WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.

By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
arXiv AI
Sep 7

Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning

The paper addresses performance discrepancy in cross-domain 3D class‑incremental learning, where 3D point clouds from heterogeneous sources cause varying degrees of performance loss beyond catastrophic forgetting. The authors introduce the Domain3D‑CIL protocol and adapt existing CIL methods to 3D, showing consistent discrepancy across baselines. They propose PolyMem, an exemplar‑free approach that models high‑order feature statistics to improve cross‑domain robustness and reduce performance discrepancy.

By Jinge Ma, Gautham Vinod, Bruce Coburn, Jui-Feng Chi, Siddeshwar Raghavan, Fengqing Zhu
arXiv Computation and Language
Sep 7

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

The paper introduces ROBORMBENCH, a benchmark comprising 2,390 real‑robot trajectories, 21,673 verified paraphrases, and ground‑truth progress labels, to evaluate paraphrase robustness in vision‑language reward models (VLMs). It demonstrates that current VLMs often give different rewards for semantically equivalent goal descriptions, sometimes flipping a robot’s outcome from failure to success. The study finds that this instability is widespread, worsens with more divergent rewrites, and is not mitigated by model scale or explicit reasoning, though dedicated reward models trained with trajectory‑grounded supervision show greater stability.

By Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No
arXiv AI
Sep 7

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

The paper introduces CABAL, an end-to-end multi-agent simulation framework that models reviewer assignment in academic conferences using large language model-driven reviewer agents. It presents an affinity-guided collusive bidding strategy that forms collusion rings based on reviewer-paper affinities, leading to more effective target-paper capture and higher scores for colluding reviewers. Experiments show that while collusive bidding significantly increases target-paper capture and reviewer scores, overall conference-wide effects are modest, and existing bid-phase detectors offer limited detection capability.

By Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li, Zhuan Shi, Fengpeng Li, Haiwei Wu, Jiantao Zhou
arXiv AI
Sep 7

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution

The paper introduces a Cost-Aware Hierarchical Multi-Agent System (HMAS) for ransomware detection and family attribution that adaptively selects analysis modalities to balance accuracy and computational cost. Static analysis is used first, with dynamic and memory modalities added only when confidence is low or specialist agents disagree, guided by a cost model. Experiments show HMAS achieves high accuracy (96.57% binary detection, 0.90 macro‑F1 attribution) while reducing analysis cost by 43.97% and latency, with 56.05% of cases resolved using static evidence alone.

By Mubashar Iqbal, Asifullah Khan
arXiv AI
Sep 7

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

MultihopSpatial is a new benchmark for Vision‑Language Models that focuses on multi‑hop, compositional spatial reasoning with queries ranging from 1 to 3 hops across varied spatial perspectives. It introduces the Acc@50IoU metric, which jointly evaluates answer selection and precise bounding‑box prediction, and provides a large‑scale training corpus, MultihopSpatial‑Train, to improve spatial intelligence. Evaluation of 37 state‑of‑the‑art VLMs shows that compositional spatial reasoning remains a significant challenge, and reinforcement learning fine‑tuning on the corpus boosts both intrinsic spatial reasoning and downstream embodied manipulation performance.

By Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
arXiv AI
Sep 7

A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

The paper presents a decentralized navigation framework for composite heterogeneous robots that integrates a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller. Each robot independently generates and refines policies at the round level using LLM inference, while the Double DQN handles tick-level action selection based on navigation variables and LLM priors. Across 30 rounds, the full configuration achieved all goals with the lowest median completion time (42 ticks) and a 25–39% improvement over other setups.

By Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif, Carlo R. daCunha
arXiv Computer Vision
Sep 7

E-RGB-D: Real-Time Event-Based Perception with Structured Light

The paper introduces E‑RGB‑D, a real‑time event‑based perception system that combines a Digital Light Processing projector with a monochrome event camera to produce RGB‑D data. By projecting structured light and capturing asynchronous brightness changes, the system can detect color and depth for each pixel, achieving a color detection speed of 1400 fps and a depth detection rate of 4 kHz. The approach enables frameless RGB‑D sensing and delivers colorful point clouds without compromising spatial resolution.

By Seyed Ehsan Marjani Bajestani, Giovanni Beltrame
arXiv AI
Sep 7

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente ’Optima is an agentic framework that builds and manages Bayesian optimization campaigns for self‑driving laboratories, separating large language model reasoning from campaign execution. It maintains a persistent optimization state, allowing consistent repetitive loops and auditable decisions, and only returns control to the agent when interpretation or revision is needed. In tests on digital discovery tasks and physical platforms, it corrected measurement failures, improved yields, and recommended formulation changes, outperforming human‑directed campaigns in cost and material usage.

By Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv AI
Sep 7

How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions

The paper reports the first empirical study comparing how humans and large language models (LLMs) evaluate perceived moral agency (PMA) in both human and autonomous artificial agents within smart city scenarios. Using a validated PMA scale, 190 human participants and various LLMs were assessed, revealing that humans are perceived to have higher moral agency than artificial agents. When confronted with moral dilemmas, LLMs focus on situational factors such as harm severity and urgency, mirroring the context‑sensitivity observed in human raters.

By Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen
arXiv AI
Sep 7

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

The paper introduces Linguistic Trajectory Encoding (LTE), a hybrid representation that compresses dynamic object motion histories using natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location while preserving accuracy with geometric waypoints and linguistic descriptions. Evaluated on the newly constructed Spatial Memory Benchmark (SMB) from EgoLife multi‑day recordings, LTE achieves 45.3 % success in semantic trajectory retrieval and 48.7 % in long‑horizon object retrieval, outperforming prior structured‑memory and VLM baselines, and compresses trajectories 8.7×–26.1× with sub‑second query latency on 24‑hour video.

By Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li, Zhan Xu, Jian Yang, Lanjun Wang, Zili Yi
arXiv AI
Sep 7

One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

The paper presents a single pretrained diffusion traffic model that serves both as an ego motion planner and as a controllable generator of safety‑critical scenarios for autonomous driving. It introduces a Single‑Stream Dual‑Stream diffusion‑transformer decoder (SSDS) that fuses scene context via joint attention, improving closed‑loop performance on the nuPlan benchmark, and a training‑free guidance scheme called Decoupled Annealing Posterior Sampling with Energy (DAPSE) that injects arbitrary energy functions at inference time. Using the same model, the authors generate realistic long‑tail driving interactions—such as aggressive cut‑ins and lead‑vehicle braking—through inference‑time guidance, exposing failure modes in black‑box planners that standard benchmarks miss.

By Arka Pal, Rajesh Kumar, Hannes Eriksson, R\'emi Lacombe, Arvid Laveno Ling, Ankit Gupta, Maciej Wozniak
arXiv AI
Sep 7

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to evaluate Vision‑Language‑Action models on fine‑grained spatial reasoning and long‑horizon procedural planning. It contains 10 task categories, 56 base tasks, and 280 variants across five difficulty levels, with 527K trajectories collected from multiple embodiments and scenes. The benchmark introduces diagnostic metrics beyond binary success, revealing that current VLA models struggle with complex spatial relations, precise execution, and memory‑intensive planning.

By Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv Machine Learning
Sep 7

A Sim-to-Real Study of Surface-Code Decoder Benchmarking

The study benchmarks six quantum error‑correction decoders on the Willow processor, the first device operating below the surface‑code threshold, using a hierarchy of increasingly realistic noise models. By evaluating real hardware data across multiple code distances, bases, and round counts, the authors find that rank agreement with hardware emerges only when each operation type is assigned its own error rate. They also independently test NVIDIA’s Ising pre‑decoder, showing it offers no accuracy‑latency advantage over other decoders in most evaluations, and release the full evaluation pipeline and data for future comparisons.

By Shay J. Manor, Leila S. Erhili, Yassine Jebbouri
arXiv Computer Vision
Sep 7

LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering

LensStyle is a unified framework for controllable stylized lens effect rendering that explicitly models lens aesthetics through joint continuous‑discrete control. It uses a Dual‑Path Controller to separate continuous optical parameter modulation (e.g., focus distance, blur strength) from discrete lens‑style conditioning (e.g., circular, polygonal, donut, cat‑eye, starburst effects), allowing fine‑grained, interpretable, and physically grounded manipulation. The authors also curate a MultiLens dataset of multi‑lens image pairs synthesized under real optical constraints, and experiments show LensStyle outperforms existing lens effect rendering methods and diffusion‑based image editing models in realism, controllability, and aesthetic quality.

By Yachuan Huang, Liwen Xiao, Liao Shen, Qiwen Wang, Huiqiang Sun, Zhiyu Pan, Zhiguo Cao
arXiv Computer Vision
Sep 7

HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction

HiSfM introduces a hierarchical coarse‑to‑fine Structure‑from‑Motion framework that enhances robustness and efficiency by building a scaffold of local communities and a compact skeleton using edge‑disjoint spanning trees. The method verifies skeletal edges with a two‑view disambiguator, constructs a stable scaffold as an anchor, and then registers remaining images for refinement. Experiments on ambiguity‑focused benchmarks and general datasets demonstrate that HiSfM avoids ambiguity‑induced failures, reduces runtime, and improves completeness compared to prior approaches.

By Ziding Zhao, Hainan Cui, Peilin Tao, Shuhan Shen
arXiv Machine Learning
Sep 7

Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics

Squint is a visual Soft Actor Critic algorithm designed to accelerate reinforcement learning for robotics. It combines parallel simulation, a distributional critic, resolution squinting, layer normalization, a tuned update-to-data ratio, and an optimized implementation to reduce wall‑clock training time. On the SO‑101 Task Set, Squint trains policies in as little as 15 minutes on a single RTX 3090 GPU, with most tasks converging in under 6 minutes and successfully transferring to a real SO‑101 robot.

By Abdulaziz Almuzairee, Henrik I. Christensen
arXiv Computer Vision
Sep 7

VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

VoxelFix is a graph‑based post‑hoc semantic correction method that refines voxel labels in completed 3D voxel maps while preserving their geometry and occupancy. It learns to correct errors by exploiting local geometry and neighboring semantic information, using training pairs generated by corrupting annotated maps with class confusions from upstream perception pipelines. Experiments on OccuFly maps show consistent improvements of 4.23–5.00 percentage points in mIoU, especially for tree, roof, and wall classes, and the method generalizes to out‑of‑distribution aerial scenes.

By Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer, Toma\v{z} Coti\v{c}, Sivasubiramaniam Subbiah, Andreas Greiner, Paul Spannaus, Sebastian Houben
arXiv Computer Vision
Sep 7

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

The paper proposes a neuro‑symbolic framework that augments vision‑language‑action (VLA) models with explicit task graphs and multimodal procedural memory to handle long‑horizon manipulation tasks. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory tracks the active step, completed actions, textual context, and relevant visual evidence. Human demonstrations provide spatial and temporal guidance via gaze or saliency cues, which are annotated in robot‑view teleoperation videos and used to fine‑tune VLA models. The approach is evaluated on workspace clearing and surgical‑instrument handling tasks, measuring object and destination selection, subtask completion, task progress, step‑order consistency, overall success, and procedural or execution mistakes.

By Vivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger