Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

4,106 stories · RSS feed

arXiv AI
Sep 10

A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing

The paper introduces a multi‑modal late‑fusion perception pipeline for object detection and tracking in autonomous racing. It combines independent detections from cameras, LiDARs, and RADARs to produce timely and robust state estimates of surrounding vehicles. The tracking framework compensates for detection delays and incorporates vehicle dynamics and track layout knowledge, and its effectiveness is confirmed through real‑world experiments in diverse critical scenarios.

By Davide Malvezzi, Michele Pestarino, Vittoria Cavicchioli, Valentina La Gamba, Silvia Severi, Fabio Bagni, Luca Bartoli, Massimiliano Bosi, Francesco Gatti, Micaela Verucchi, Ayoub Raji, Marko Bertogna
arXiv AI
Sep 10

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.

By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
arXiv Machine Learning
Sep 10

ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages

ComplicitSplat is a novel black‑box attack that leverages 3D Gaussian Splatting (3DGS) shading to create viewpoint‑specific camouflage, embedding adversarial content into scene objects that is only visible from certain angles. The method does not require access to model architecture or weights and can successfully fool a range of popular object detectors—including single‑stage, multi‑stage, and transformer‑based models—on both real‑world physical objects and synthetic scenes. This demonstrates that downstream models using 3DGS are vulnerable to adversarial manipulation.

By Matthew Hull, Haoyang Yang, Pratham Mehta, Mansi Phute, Aeree Cho, Haorang Wang, Matthew Lau, Wenke Lee, Wilian Lunardi, Martin Andreoni, Duen Horng Chau
arXiv AI
Sep 10

WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting

WildRelight is the first in-the-wild dataset designed to evaluate single-image relighting models, featuring high-resolution outdoor scenes captured under strictly aligned, temporally varying natural illuminations paired with high-dynamic-range environment maps. The benchmark demonstrates that state-of-the-art models trained on synthetic data suffer severe domain shifts when applied to real-world imagery. Leveraging the dataset’s temporal structure, the authors introduce a physics-guided inference framework combining Diffusion Posterior Sampling with Temporal Sampling-Aware Test-Time Adaptation, enabling synthetic models to self-supervise and align with real-world statistics on-the-fly.

By Lezhong Wang, Mehmet Onurcan Kaya, Siavash Bigdeli, Jeppe Revall Frisvad
arXiv Machine Learning
Sep 10

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

arXiv:2609.08084v1 Announce Type: cross Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...

By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
arXiv Machine Learning
Sep 10

BinauralVAE: Spatial Audio Reconstruction For World Models

BinauralVAE is an open‑source pipeline that reconstructs spatial audio using various Variational Autoencoder architectures, including complex‑valued variants, to learn latent representations of binaural signals. The project builds on realistic acoustic data from a simulated robot navigating an environment, providing a foundation for audio‑centric world models. It aims to map the causal link between navigational actions and their acoustic outcomes, positioning sound as a complementary modality for spatial awareness.

By Luis Vitor Zerkowski, Luiz Velho
arXiv AI
Sep 10

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

AtomicVLA is a unified planning-and-execution framework that generates task-level plans, atomic skill abstractions, and fine-grained actions for robotic manipulation. It builds a scalable atomic skill library using a Skill‑Guided Mixture‑of‑Experts (SG‑MoE) and a flexible routing encoder that assigns new skills to dedicated experts, enabling continual learning. Experiments show that AtomicVLA outperforms baseline models on both simulated and real‑world long‑horizon tasks, achieving significant improvements in task performance and learning efficiency.

By Likui Zhang, Tao Tang, Zhihao Zhan, Xiuwei Chen, Zisheng Chen, Jianhua Han, Jiangtong Zhu, Pei Xu, Hang Xu, Hefeng Wu, Liang Lin, Xiaodan Liang
Hugging Face Trending Papers
Sep 8

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

SyncWorld is an action‑conditioned world model that functions as a zero‑shot simulator across unseen environments without additional training. It uses a visual calibration episode—paired frames and actions that expose all controllable degrees of freedom—to define a setup‑specific Action‑Visual Mapping in context. By training with these calibration contexts, the model learns to interpret actions through visual evidence and to leverage interaction history when explicit calibration is unavailable, enabling accurate simulation of action outcomes and test‑time policy improvement.

Hugging Face Trending Papers
Sep 8

A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes

The paper presents a four‑module, data‑driven framework to identify and prioritize robotic process automation (RPA) opportunities in U.S. hospitals. It includes a process taxonomy, an automation suitability index, a tool‑tier selection recommendation, and a return‑on‑investment analysis, all applied to a synthetic portfolio of twenty hospital processes. The authors demonstrate the framework’s robustness through Monte Carlo simulations and discuss governance and future validation steps.