Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

4,108 stories · RSS feed

arXiv Machine Learning
Sep 1

What Emerges and What Breaks in Self-Play Driving

The paper reports on training autonomous driving policies via self‑play, extending previous work by using Transformers and a real‑city high‑definition map. On CARLA and Waymo benchmarks, the resulting policies underperform compared to Gigaflow, with identified failure modes such as reward hacking at traffic lights and lack of incentive to stop at stop signs. The authors also analyze which traffic rules emerge from self‑play and confirm that reward conditioning produces diverse driving behaviors.

By Laur Sisask, Ardi Tampuu, Tambet Matiisen
arXiv AI
Sep 1

Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers

The paper introduces a fully distributed continuous‑time algorithm for solving Generalized Nash Equilibrium Problems (GNEPs) with shared linear equality constraints. Unlike existing methods that require exchanging Lagrange multipliers, this approach converges to any GNE without multiplier communication, thereby reducing communication overhead and enhancing privacy. Discrete‑time variants are also presented and the method is demonstrated on a multi‑robot placement task.

By Shao-An Yin, Mingyi Hong, Nicola Elia
arXiv AI
Sep 1

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.

By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
arXiv Computation and Language
Sep 1

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.

By Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh
arXiv Computer Vision
Sep 1

Proximity3D: Shape from Capacitive Proximity on Sensing Manifold

arXiv:2608.30344v1 Announce Type: new Abstract: Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images or depth maps. In this paper, we use a cu...

By Hao Chen, Chenming Wu, Chun Ping Lam, Xiangjia Chen, Guoxin Fang, Charlie C. L. Wang, Yeung Yam, Juncong Lin, Chengkai Dai
arXiv AI
Sep 1

Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration

The paper introduces OHCAM, an online method for learning action models that include conditional and quantified effects from limited interactions. It maintains a belief over possible models and actively chooses actions that maximize disagreement among hypotheses to reduce uncertainty, while handling noisy observations. Starting with simple hypotheses, OHCAM expands complexity only when necessary, achieving sample‑efficient learning that outperforms baselines on benchmark domains and is validated on a Kinova Gen3 robot.

By Jeffrey Jewett, William Solow, Sandhya Saisubramanian