Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,858 stories · RSS feed

arXiv Computer Vision
Sep 17

StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions

arXiv:2609.18430v1 Announce Type: new Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...

By WM Team, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu
arXiv Computer Vision
Sep 17

In-Context Robot Learning with VLM Agents

arXiv:2609.19138v1 Announce Type: new Abstract: Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations...

By Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu
arXiv AI
Sep 17

WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories

WetRobo is a reproducible robot kit designed to enable wet‑lab researchers to delegate tasks to coding agents without teleoperation or neural‑network training. The kit includes a robot arm, essential lab equipment, pre‑recorded teleoperation demos, and a skill file, allowing a coding agent to observe the lab, write, and execute programs using external tools. Experiments with OpenAI Codex on tasks such as lifting a Petri dish lid, removing a bottle cap, and opening an incubator door demonstrated successful performance in two different laboratories, outperforming a fine‑tuned vision‑language‑action policy that failed to transfer.

By Yuna Oikawa, Kei Endo, Takanori Uzawa, Yunzhe Zhang, Manan Anjaria, Lerrel Pinto, Sherry Yang, Koji Tsuda
arXiv AI
Sep 17

REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

REVERSAL-BENCH is a benchmark that introduces a continuous reversibility parameter ρ∈[0,1] and a reset oracle to evaluate how well reinforcement learning agents can recover from irreversible states across eight manipulation tasks in five physics engines. Experiments show a sharp reversibility cliff: reset‑free agents become trapped in irrecoverable states as ρ increases, while episodic agents continue learning steadily. The benchmark also provides a large multi‑simulator dataset and demonstrates that safety shields can predict recoverability but only succeed when the agent can avoid the trap.

By Riyaaz Shaik, Chandru Venkataraman
arXiv AI
Sep 17

A Comprehensive Review of Generative Physical Artificial Intelligence

The paper surveys Generative Physical Artificial Intelligence (GPAI), a field where large foundation models are integrated with physical robots. It introduces a taxonomy of five approaches—Robot Foundation Models, Vision‑Language Action models, Large Behavior Models, Diffusion Policy Models, and World Foundation Models—and discusses how they complement each other across domains such as autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems. The review highlights performance gains, data‑efficient learning, sim‑to‑real transfer, edge‑compatible architectures, and safety frameworks as key research directions.

By Satyam Gaba, Krutiksinh Rana, Siva Sai, Vinay Chamola, Dusit Niyato
arXiv Machine Learning
Sep 17

Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL

Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.

By Everest Yang
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Sep 17

Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception

The paper examines whether standard image‑fidelity metrics (SSIM, PSNR, MSE) accurately reflect the performance of GAN‑generated synthetic sonar data in robotic perception tasks. Using a Pix2Pix GAN with four discriminator configurations (PixelGAN, PatchGAN‑16, PatchGAN‑70, ImageGAN), the authors train object detectors (YOLOX‑S, YOLOX‑L, Faster R‑CNN) solely on real sonar images and evaluate them on the synthetic outputs. Results show a mismatch: the discriminator that yields the best pixel‑level scores does not always produce the best detection performance, with PatchGAN models achieving strong downstream results despite lower SSIM/PSNR/MSE values.

By Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns
arXiv AI
Sep 17

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

The paper introduces CSWAM, a Causal Semantic World Action Model that enhances FastWAM by integrating a causal semantic expert based on V-JEPA 2.1. This expert provides temporally grounded, appearance‑agnostic representations of semantic state changes and motion, leveraging sparse observation history and causal attention to improve action‑only inference. Experiments on simulation and real‑robot tasks show that CSWAM significantly boosts out‑of‑distribution generalization, raising success rates from 10.16% to 45.18% on RoboTwin 2.0 and from 27.5% to 70.0% across real‑robot tasks.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv AI
Sep 17

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

The paper introduces a pipeline that combines generated video and audio to produce force-aware manipulation trajectories for a Franka Panda robot. By using the loudness of contact sounds to shape a bounded, time-varying desired-force profile, the system can execute tasks that require precise contact forces, outperforming kinematic-only baselines. The approach also serves as a data generation engine for training closed-loop policies.

By Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
arXiv Machine Learning
Sep 17

RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control

RecMorph introduces a topology‑guided spatial recurrent architecture for generalized morphology control, converting a kinematic tree into a sequence that enables joint cross‑limb communication and representation transformation. The design incorporates residual preservation, RMS normalization, and input‑dependent channel modulation to stabilize repeated spatial transformations, achieving linear token complexity. Across five UNIMAL tasks and a four‑platform quadruped setting, RecMorph outperforms existing controllers in training performance, inference throughput, and generalization to unseen bodies with up to 30 limbs, while also demonstrating robust real‑world performance on Go1/Go2 trials.

By Quanrui Rao, Yong Liu, Xueming Xiao, Yingbo Luo, Kun Wu, Zhenyu Xu, Meibao Yao
arXiv AI
Sep 17

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.

By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu