Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv Computer Vision
4d ago

Resolving Mixed Single-Photon LiDAR Returns for Foreground-View and Hidden Scene Reconstruction

The paper introduces a state‑aware framework that reconstructs both foreground and hidden scenes from single‑photon LiDAR histograms affected by partially transmissive occluders. It classifies each ray into no‑return, single‑return, or dual‑return states, guiding a two‑head neural field that jointly refines waveform reconstruction and geometry localization. A new paired dataset of occluded and clean LiDAR captures validates the method, showing improved depth and point‑cloud accuracy over existing baselines.

By Ziting Wen, Runrong Deng, Zili Zhang, Haitao Zheng, Yuecong Xu, Xiaoqiang Ren, Guodong Shi, Kemi Ding
arXiv AI
4d ago

Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots

The paper investigates the gap between intent and behavior in jailbreak attacks on LLM-based robots, showing that many intent jailbreaks fail to produce harmful physical actions due to robot-specific constraints. It introduces POEF, an automated red‑teaming framework that optimizes jailbreak prompts using hidden‑layer gradients and evaluates policy feasibility with a multi‑agent system, achieving an 80% success rate on commercial robots. The authors also propose defense strategies and emphasize the need for stronger countermeasures before widespread deployment of LLM-based robots.

By Xuancun Lu, Zhengxian Huang, Xinfeng Li, Chi Zhang, Xiaoyu Ji, Wenyuan Xu
arXiv Computer Vision
4d ago

3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability

The paper introduces 3DROID, a dataset of renderable 3D Gaussian scenes that are anchored to a robot’s metric workspace and include per-scene reliability metrics. It examines how the reliability of camera extrinsics and pose conditioning affect the fidelity of 3D Gaussian representations, proposing a calibration-aware pipeline that improves novel-view rendering when extrinsics are trustworthy. The resulting dataset, available on Hugging Face, provides robot manipulation researchers with metric-scale, pose-anchored 3D data and reliability annotations.

By Wonguen Cho, Junhoo Lee, Nojun Kwak
arXiv Computer Vision
4d ago

UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

UniTrackPLA introduces a unified panorama-language-action model that simultaneously handles instruction‑guided navigation and dynamic person tracking for embodied robots. Its Panoramic‑Aware Encoding preserves azimuthal and temporal structure, allowing a shared vision‑language backbone to generate continuous waypoint chunks for both tasks. The model also employs World‑Action Consistency to predict future visual states and verify waypoint prefixes, enabling reliable action reuse and replanning when inconsistencies arise. A new OmniTrackNav‑Bench dataset and extensive real‑world experiments demonstrate significant performance gains over prior methods.

By Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang
arXiv Computer Vision
4d ago

Beyond Leaderboard Scores: A Deployment-Focused Protocol for Interpretable Tracking Evaluation in Pedestrian-Centric Environments

The paper introduces a deployment‑focused evaluation protocol for pedestrian tracking that isolates tracker performance by using shared detections, assessing initialization, gap continuation, identity recovery, close‑neighbor association, and runtime under load, while still reporting HOTA as an aggregate metric. Applied to the JackRabbot Dataset, the protocol reveals that trackers vary widely in capability, with none maintaining spatially correct, same‑identity output after 1 s of missing detector support, and that runtime on an NVIDIA Jetson Orin often falls below the 10 Hz target in crowded scenes. The authors provide open‑source code and scripts to enable reproducible assessment of tracker suitability for real‑world pedestrian‑centric robotics. whyItMatters":"The protocol offers a practical, tracker‑only framework that directly informs the design and deployment of mobile robots operating among pedestrians by highlighting specific failure modes and computational constraints that aggregate scores overlook."

By Dominik Wojcikiewicz, Diego Paez-Granados
arXiv Computer Vision
4d ago

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

LangMap introduces a human‑verified benchmark for language‑conditioned goal navigation (LGN) that spans four hierarchical semantic levels—scene, room, region, and instance—within real‑world indoor 3D scans. The dataset, built on HM3D, contains 18K tasks with concise and detailed descriptions for 414 object categories, and its contrastive annotation protocol ensures high‑quality, discriminative region and instance labels. Evaluation shows that LangMap’s descriptions improve text‑to‑view matching accuracy by 23 points over GOAT‑Bench and achieve a 92.5% unique‑and‑correct match rate in an independent human audit, while a proposed RGB‑only baseline, PlaNaVid, attains top‑tier success rates without depth or 3D scene representations.

By Bo Miao, Weijia Liu, Jun Luo, Lachlan Shinnick, Jian Liu, Thomas Hamilton-Smith, Yuhe Yang, Zijie Wu, Vanja Videnovic, Feras Dayoub, Anton van den Hengel
arXiv Computer Vision
4d ago

AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning

The paper introduces AEGIS, a buffer‑free, layer‑wise orthogonal gradient projection framework that enables continuous fine‑tuning of Vision‑Language Models for robotic manipulation while protecting pre‑trained visual reasoning. AEGIS uses pre‑training activation statistics as static anchors and applies a Wasserstein‑2 transport penalty to generate an anchor‑restoration gradient, followed by a dual‑backward pass that orthogonalizes task gradients against this restoration vector. Experiments on the LIBERO manipulation benchmark with PaliGemma2‑3B show that AEGIS preserves Visual Question Answering performance and baseline loss while achieving continuous action convergence, all without replay buffers, teacher models, or co‑training data.

By Guransh Singh
arXiv AI
4d ago

Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking

The paper introduces Latent Frequency Masking, an attack that removes or weakens invisible watermarks in AI-generated images by altering selected Fourier coefficients in the latent space. It demonstrates that this method can effectively compromise six diffusion watermarking techniques while maintaining image quality and running faster than existing attacks. The study provides a theoretical distortion bound and evaluates the attack on images from DiffusionDB and MS-COCO prompts.

By Kirill Aistov, Khaled Abud, Irina Serzhenko, Egor Kovalev, Aleksey Yakushev, Aleksandr Akimenkov, Dmitry Obydenkov, Yury Markin, Sergey Lavrushkin, Dmitriy Vatolin, Anastasia Antsiferova
arXiv AI
4d ago

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

HumanoidToolBench is an 18‑task benchmark designed to evaluate humanoid robots on tool selection, manipulation, and locomotion across three scenarios, execution levels, and tool‑set modes. It is accompanied by ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1 robot. Experiments with seven policies in simulation and three on the real robot show significant gaps between tool selection and task completion, especially on unseen tools and under unrelated instructions.

By Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, Idan Szpektor, Avi Caciularu, Jongmin Park, Youngjae Yu
arXiv AI
4d ago

Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination

The paper introduces Watch, Infer, Coordinate, a benchmark and inference method for robots to deduce a partner’s physical constraints by observing joint behavior in physically coupled manipulation tasks. It demonstrates that understanding these constraints enables zero‑shot coordination on new tasks, achieving performance close to an oracle with true constraint knowledge. The study focuses on scenarios where a robot’s limitations may be unknown to its partner, such as hardware degradation or actuator faults.

By Suyu Ye, Zheyuan Zhang, Vaishnav Tadiparthi, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Tianmin Shu, Homanga Bharadhwaj, Nakul Agarwal
arXiv AI
4d ago

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Reconstruct, Practice, Go Real (RPG) is a framework that enables autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities from offline data, creates simulation practice tasks, and uses feedback to diagnose failures, develop new symbolic skills, refine existing ones, and revise the system prompt. Across 22 manipulation tasks, RPG raises task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming baselines and achieving perfect success on 30 physical trials after calibration.

By Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
arXiv AI
4d ago

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.

By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
arXiv AI
4d ago

Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities

The paper introduces Token Communication (TokCom) as a native interface for Collaborative Embodied Artificial Intelligence (CEAI), where tokens act as compact semantic carriers and inference units for generative foundation models. It outlines a TokCom-assisted CEAI framework featuring a task‑adaptive communication protocol that includes a compact codebook, syntax rules, and contextual examples to guide token distillation and reconstruction over wireless channels. A case study on collaborative object transport shows that TokCom significantly reduces source payload bit consumption while maintaining task efficiency and robustness to noisy channels.

By Peng Yi, Ying-Chang Liang
arXiv AI
4d ago

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

The paper introduces 3D-Prog, a framework that adapts powerful 2D vision‑language models (VLMs) for reliable 3D understanding, manipulation, and generation. It does so by employing Canonical Coordinate Framing (CCF) to anchor inputs and outputs in a shared Euclidean coordinate system, resolving axis ambiguity, metric scale, and reference issues, and Task‑Adaptive Feedback (TAF) to close the reasoning loop with dynamic, task‑specific feedback. Together, CCF and TAF enable 2D VLMs to perform open‑vocabulary 3D tasks without retraining, yielding consistent, interpretable, and high‑quality results across diverse 3D scenarios.

By Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel
arXiv AI
4d ago

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

DuoMind is a distributed hierarchical framework that enables multi-robot coordination by combining vision-language-action models for low‑level execution with vision-language models for high‑level reasoning and inter‑robot communication. Each robot’s orchestrator processes task instructions, local observations, and messages from peers to generate precise action commands and semantic messages for other robots. The authors introduce RoboPoly, a benchmark of long‑horizon manipulation tasks, and show through experiments that DuoMind improves multi‑robot task performance, with ablation studies highlighting the roles of hierarchical orchestration and semantic communication.

By Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
arXiv AI
4d ago

Probing an Embodied LLM: When Higher Observation Fidelity Hurts Problem Solving

The study investigates how different observation fidelities affect the problem‑solving performance of embodied large language model (LLM) agents. Using the Lockbox mechanical puzzle, agents were tested with raw RGB, RGB‑D, and perfect ground‑truth symbolic observations in both physical and simulated environments. Surprisingly, agents performed best with raw RGB input and worst with perfect observations; introducing moderate perceptual noise in simulation further improved success rates, suggesting that some errors can mitigate repetitive action loops.

By Oussama Zenkri, Oliver Brock
arXiv AI
4d ago

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

UrbanVLA is a Vision‑Language‑Action framework designed to enable delivery robots to navigate large‑scale urban environments using long‑horizon route instructions. The model aligns noisy route waypoints with visual observations and plans trajectories, trained through a two‑stage pipeline of supervised fine‑tuning on simulated data and reinforcement fine‑tuning on mixed simulation and real‑world data. Experiments show UrbanVLA outperforms strong baselines by over 55% on the SocialNav task and demonstrates reliable real‑world navigation in large urban settings.

By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang
arXiv AI
4d ago

Constant-Time Planning for Chaining Collision-free Motion to Manipulation Behaviors

The paper introduces the Behavioral Constant-Time Motion Planner (B-CTMP), an extension of Constant-Time Motion Planning that handles two-step manipulation tasks in semi-structured environments. B-CTMP constructs neighborhoods in object-pose space and uses statistical certification to ensure a user-specified success rate, caching plans only when repeated rollouts meet this threshold. The method is evaluated on shelf picking, plug insertion, and wheel replacement, showing consistent success where baseline planners fail and rejecting infeasible poses in constant time.

By Nayesha Gandotra, Itamar Mishani, Lai Yuan, Oren Salzman, Maxim Likhachev
arXiv AI
4d ago

Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

Rewind-IL is a training‑free online safeguard for generative action‑chunked imitation learning policies. It uses a zero‑shot failure detector based on Temporal Inter‑chunk Discrepancy Estimate (TIDE) and a state‑respawning mechanism that returns the robot to a verified safe intermediate state. The system builds a checkpoint library offline with a vision‑language model and monitors self‑consistency online, rewinding execution to the latest safe checkpoint when a failure is detected, thereby improving reliability in long‑horizon manipulation tasks.

By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi