Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,858 stories · RSS feed

arXiv AI
Sep 21

Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments

arXiv:2609.21828v1 Announce Type: cross Abstract: Blind and low-vision users often face challenges when locating and physically acquiring objects in unfamiliar indoor environments. Existing vision-la...

By George Xi Wang, Xiangyu Li, Shaoyue Wen, Jiaqian Hu, Junan Xie, Yupeng Wang, Ziyue Shi, Qijun Chen, Maaike Bouwmeester, Yuhua Jin, Jing Qian
arXiv AI
Sep 21

Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning

arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...

By Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong
arXiv Computer Vision
Sep 21

Adaptive Color Grading

arXiv:2609.21169v1 Announce Type: cross Abstract: Independent control of tonescale regions (e.g., shadows, highlights) is essential for painters, photographers and cinematographers to bring 2D images...

By Trevor D. Canham, Abhijith Punnappurath, Michael S. Brown
arXiv Computer Vision
Sep 21

Refining Ground Truth Poses in Autonomous Driving Datasets via Neural Rendering

arXiv:2504.15776v2 Announce Type: replace Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...

By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv Computer Vision
Sep 18

BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

BinoGen is an automated framework that generates large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. It models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups, producing synchronized videos with dense multimodal supervision such as depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, the authors created a dataset of over 20 million annotated images, demonstrating that incorporating this data improves real-world visual perception tasks like depth estimation, object detection, and video object tracking, and that embodiment-specific adaptation enhances performance while joint training enables a single model to perform competitively across different observer embodiments.

By Chunpeng Li, Ya-tang Li
arXiv AI
Sep 18

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

SIMLIFE is a scalable platform that simulates long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues. It introduces the SimLife-BP benchmark, which tests long-context pattern understanding by requiring agents to infer latent behavioral rules from weeks or months of everyday observations across 106 episodes. The benchmark includes 1,439 question-answer pairs that probe direct, counterfactual, noisy, and inverse reasoning under varying rule hints.

By Run Peng, Zinnia Nie, Jing Ding, Yinpei Dai, Yichi Zhang, Zengqing Wu, Yao Fu, Ziqiao Ma, Jiayuan Mao, Joyce Chai
arXiv AI
Sep 18

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.

By Albert Wu, Nicholas Roberts, Tzu-Heng Huang, Haoran Lin, Gil Friedman, Sungjun Cho, Gabriel Orlanski, Frederic Sala
arXiv AI
Sep 18

MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution

MAGMA-GEN is an on‑policy data‑generation pipeline that transforms ambiguous failures in hierarchical robotic manipulation into validated recovery supervision. It uses a privileged coach to hypothesize early decision‑level errors and proposes localized corrections, then retains only those candidates that improve downstream progress when re‑executed from the same state. This approach generates supervised examples from the agent’s own failure distribution, enabling improved task success and recovery without requiring per‑step human demonstrations.

By Loan Bernat (LAAS-GEPETTO), Matthieu Grard (LAAS-RAP), Ariane Herbulot (LAAS-RAP), Florent Lamiraux (LAAS-GEPETTO)
arXiv AI
Sep 18

JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

JEPA-WAM enhances World Action Models (WAMs) by pairing text instructions with stochastically generated visual cues, using a text-to-image generator and a frozen V‑JEPA encoder to create dense goal representations. These representations are compressed into goal tokens that condition both video and action experts via cross‑attention, enabling the model to better ground instructions. On a new real‑robot benchmark, JEPA‑WAM attains 87.3%, 74.5%, and 80.9% success rates across in‑distribution, out‑of‑distribution scenes, and out‑of‑distribution instructions, outperforming prior methods by significant margins.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv AI
Sep 18

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.

By Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius, Daniel J. Szafir, Rudolf Lioutikov