Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,943 stories · RSS feed

arXiv Machine Learning
Sep 17

Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL

Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.

By Everest Yang
arXiv AI
Sep 17

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

The paper introduces CSWAM, a Causal Semantic World Action Model that enhances FastWAM by integrating a causal semantic expert based on V-JEPA 2.1. This expert provides temporally grounded, appearance‑agnostic representations of semantic state changes and motion, leveraging sparse observation history and causal attention to improve action‑only inference. Experiments on simulation and real‑robot tasks show that CSWAM significantly boosts out‑of‑distribution generalization, raising success rates from 10.16% to 45.18% on RoboTwin 2.0 and from 27.5% to 70.0% across real‑robot tasks.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv AI
Sep 16

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

HumanEgo is a framework that enables zero‑shot robot learning from short egocentric human videos by converting each demonstration into an entity‑level hand‑object interaction representation and training a flow‑matching policy with dense auxiliary objectives. The method is robot‑data‑free, hardware‑agnostic, and data‑efficient, achieving 92.5 % success on four real‑world tasks with only 30 minutes of human video per task and outperforming matched‑time robot teleoperation by 41 %. HumanEgo also robustly transfers zero‑shot across new robots, cameras, and environments, and is released as an open‑source tool for learning robot policies directly from human data.

By Zhi Wang, Botao He, Kelin Yu, Seungjae Lee, Ruohan Gao, Furong Huang, Yiannis Aloimonos
arXiv AI
Sep 16

QDTraj: Exploration of Diverse Trajectory Primitives for Articulated Objects Robotic Manipulation

The paper introduces QDTraj, a method that uses Quality‑Diversity algorithms to automatically generate a diverse set of low‑level trajectory primitives for manipulating articulated objects. By leveraging sparse reward exploration, QDTraj produces at least five times more diverse trajectories for hinge and slider tasks compared to baseline methods, and demonstrates strong generalization across 30 articulations from the PartNetMobility dataset, averaging 704 trajectories per task. The resulting primitives are validated both in simulation and on real robots, with the code released publicly.

By Mathilde Kappel, Mahdi Khoramshahi, Louis Annabi, Faiz Ben Amar, St\'ephane Doncieux
arXiv AI
Sep 16

ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

ProxiDex is a dynamics‑guided proximity policy framework for multi‑finger dexterous manipulation that treats hand‑object proximity as an interaction state. It reconstructs interaction point clouds, converts geometric distances into proximity cues, and learns action‑conditioned proximity dynamics using a coupled forward‑inverse design. The framework adaptively reweights proximity tokens across manipulation phases and employs dynamics‑consistency supervision to stabilize action generation, leading to improved success rates and robustness in both simulation and real‑world experiments.

By Yushan Bai, Boyu Zheng, Zhiyang Mao, Hongzheng Sun, Yuchuang Tong, En Li, Zhengtao Zhang
arXiv AI
Sep 16

Bridging Learned Visual Perception and Symbolic Belief-Space Planning

The paper introduces a new paradigm called VLM-as-probabilistic-grounder, which models the uncertainty of Vision‑Language Model (VLM) predicate groundings as a probability distribution over symbolic states. This probabilistic grounding allows belief‑space planning, producing more robust plans in partially observable settings. Experiments in simulated household robot environments demonstrate that this approach improves robustness and task success compared to deterministic grounding methods.

By Guy Azran, Michael Navat, Sarah Keren