Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
Sep 28

ORCA: Evaluating LLMs on Data Science Code Translation

ORCA is a new benchmark for evaluating large language models on Data Science Code Translation (DSCT), comprising two settings: ORCA-MAIN with 1,600 grounding-level tasks across data querying, manipulation, and deep learning, and ORCA-PROJECT with 200 full-project translation tasks across seven data‑science task types. Each task includes reference translations and test cases to verify functional equivalence, and a multi‑stage quality verification process ensures task correctness. Experiments show that even state‑of‑the‑art LLMs perform poorly on DSCT, with Claude‑Opus‑4.6 achieving only 56.92% success on ORCA‑MAIN and 33.67% on ORCA‑PROJECT, while an intent‑augmented approach improves success rates by 4.80% and 5.33% respectively.

By Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng
arXiv AI
Sep 28

Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language

The paper introduces CLUE, a framework that lets robots actively resolve contextual uncertainty for underspecified natural language tasks. CLUE employs an LLM-derived policy to generate task-relevant hypotheses and plans, then uses an online language-embedded map to ground these into actions, refining its plan through closed-loop interaction. Experiments on a Boston Dynamics Spot across diverse indoor and outdoor settings show CLUE achieving near-oracle performance and outperforming LLM planners without closed-loop feedback by a significant margin.

By Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar
arXiv AI
Sep 28

Auditing Latent-Space Monitors for Autonomous Driving

The paper audits runtime failure monitors that use a model’s internal representations to predict failures in autonomous driving tasks. Across two tasks—online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD—the authors find that frame‑level errors can be predicted with high AUROC scores using supervised latent probes. However, adding latent features to baseline monitors that use only observable inputs and outputs does not yield statistically significant improvements, suggesting that internal representations may not provide additional predictive value beyond what is already observable.

By Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
arXiv AI
Sep 28

Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots

The paper introduces a design optimization framework for surgical continuum robots that jointly considers dexterity and anatomical constraints. It defines the Reachable Volumetric Dexterous Solid Angle (RVDSA) as an objective metric, uses a computationally efficient motion planner to evaluate it, and applies an asymptotically optimal simulated annealing optimizer to find improved designs. Applying the method to a bimanual dexterous sheath robot for colon procedures yielded a 78% increase in RVDSA compared to optimization based solely on 3D voxel coverage.

By Tony Qin, Peter Connor, Khoa Dang, Carter Hatch, Caleb Rucker, Robert J. Webster III, Ron Alterovitz
arXiv AI
Sep 28

NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation

NavGen introduces a text-to-video data generation pipeline that creates about 400K vision‑language navigation episodes for both indoor and outdoor scenes, using high‑fidelity visual generative models. The approach includes a style‑diversification method to scale up rare, hard‑to‑collect data. Models trained on NavGen data outperform those trained on existing UAV navigation datasets and achieve a 75% success rate in real‑world flying experiments.

By Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu, Yijin Wang, Zhiyang Liu, Fei Gao, Yuze Wu, Xin Zhou
arXiv AI
Sep 28

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.

By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
arXiv AI
Sep 28

Evaluation Is All You Need for Multi-Modal Autonomous Driving

The paper introduces iDriveVLA, a multi‑modal planning framework for autonomous driving that addresses a generation‑evaluation asymmetry by improving candidate trajectory spaces and providing a unified, safety‑aware evaluator. It combines a Safety‑aware Scorer for risk estimation with a VLM‑guided Modulator that adapts weighting to the scene, and employs an oracle‑aligned progressive training strategy. On the NAVSIM v1 leaderboard, iDriveVLA achieves a new state‑of‑the‑art PDMS score of 94.95, surpassing human‑expert performance.

By Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li
arXiv AI
Sep 28

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

MVVBench is a new benchmark for multi‑view video reasoning that tests vision‑language models on tasks requiring integration of spatial and temporal evidence across multiple, often non‑overlapping camera streams. The benchmark contains questions that cannot be answered from any single view or single moment, forcing models to jointly reason across views and time. It evaluates six capabilities—including attribute identification, relative distance, camera pose, and compositional counting—and provides human‑authored QA, rigorous verification, and detailed error analysis. "whyItMatters":"The benchmark offers a rigorous evaluation of 4D multi‑view reasoning and a foundation for future progress toward reliable embodied perception."

By Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim
arXiv AI
Sep 28

The Linear Representation Hypothesis for Vision-Language-Action Models

The paper proposes a theoretical framework called the Linear Representation Hypothesis (LRH) for vision‑language‑action (VLA) models, extending the concept from large language models to systems where physical quantities of interest (QoI) evolve with the dynamics. It introduces a signature‑based formulation that unifies representations and policies, proving that future QoI evolution can be linearly probed from representations and that a generalized linear model for stochastic action chunks allows monotonic steering of QoI. The authors validate their theory with an explicit oracle representation in a planar control‑affine navigation experiment, demonstrating the predicted linear probing and steering mechanisms.

By Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
arXiv AI
Sep 28

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

The paper proposes a new architecture for Vision‑Language‑Action (VLA) models that improves sample efficiency by training a predictive world model on the vision encoder’s embedding space. It argues that these embeddings are action‑relevant and can be used to predict future states, addressing the lack of an explicit world model in current VLAs. The trained model can also support short‑term planning by sampling actions that lead to desired goal images.

By Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter
arXiv AI
Sep 28

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.

By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
arXiv AI
Sep 28

Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers

The paper "Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers" reports that vision‑language models (VLMs) struggle to read images containing two overlapping text layers—one with sharp contour lines and one with soft shading. Using the DecoyBench dataset of 300 such images, the authors evaluated six closed‑source VLMs under naive and guided prompting at high and low resolutions. While humans could read both layers accurately, the models reliably read only the contour layer at high resolution and failed to extract the shading layer; at low resolution, neither the models nor humans could read the contour layer, but the shading layer remained readable. "whyItMatters":"The study highlights a consistent limitation of current VLMs in handling typographic structures with multiple spatial frequency layers, underscoring their vulnerability to typographic attacks and the need for more robust text‑recognition capabilities."

By Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an
arXiv AI
Sep 28

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench is a benchmark that evaluates AI safety risks in multi-agent settings, covering 1,535 high-stakes scenarios based on game-theoretic structures like the Prisoner's Dilemma, Stag Hunt, and Chicken. The benchmark draws scenarios from realistic AI risk contexts in the MIT AI Risk Repository and tests 15 frontier models, finding that agents fail to choose socially beneficial actions in 38% of cases, including military escalation, election manipulation, and medical malpractice. The study also measures how prompt framing and ordering affect outcomes and shows that game-theoretic interventions can improve socially beneficial outcomes by up to 18%.

By Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin
arXiv AI
Sep 28

Provable Speech Attributes Conversion via Latent Independence

The paper introduces a formal framework for speech attribute conversion, focusing on deterministic autoencoders with an independence constraint between latent representations and controllable attributes. It provides theoretical guarantees linking reconstruction, independence, and successful attribute manipulation under explicit population-level assumptions. The authors also propose a practical voice conversion method based on these principles and demonstrate competitive performance on voice and pitch conversion tasks.

By Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham
arXiv AI
Sep 28

Topology-Driven Anti-Entanglement Control for Soft Robots

The paper introduces a topology-driven Multi-Agent Reinforcement Learning (TD-MARL) framework designed to coordinate soft robots in precision manufacturing tasks, specifically to prevent entanglement during unwinding operations in highly constrained environments. By employing centralized learning with a shared topological state, the approach improves observability and training stability, while distributed execution reduces communication demands and enhances system reliability. Simulation results demonstrate that TD-MARL outperforms current deep reinforcement learning methods in convergence speed and anti-winding effectiveness.

By Haoyang Le, Shengxuan Wang, Mohan Chen, Shuo Feng
arXiv AI
Sep 28

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.

By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv AI
Sep 28

SUN: Agentic Robot Policy Learning with Persistent Task Programs

The paper introduces SUN (Semantically UNified) Programs, typed executables that translate grounded relations into optimal control objectives, satisfaction predicates, and learning rewards. Using the Kuafu harness, a foundation model orchestrates scene preparation, verification, residual reinforcement learning, and data generation, repairing candidate programs and calibrating reward weights. Across nine multi‑stage manipulation tasks, Kuafu achieves an 82.03% success rate, outperforms learned baselines, generates demonstrations 10.57× faster than human teleoperation, and transfers zero‑shot to physical Franka and Kinova robots.

By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong