Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
4d ago

Knowing When to Yield: Grounded Arbitration of User Corrections in Text-Based Embodied Agents

The paper introduces GAVA, a grounded arbitration framework that enables text-based embodied agents to decide whether to accept, reject, inspect, or ask for clarification when a user’s correction might be incorrect. In the ALFWorld environment, GAVA achieves perfect correction accuracy through local inspections and reduces interaction costs compared to always-verify baselines, especially when leveraging semantic priors. The study demonstrates that selective information gathering can lower declared joint costs, though it does not conclusively prove a general advantage of environmental value of information over clarification.

By Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin
arXiv AI
4d ago

Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving

The paper introduces a meta-multi-agent reinforcement learning (meta‑MARL) framework that enables rapid adaptation of interactive policies in multi‑agent systems. By modeling multi‑agent reinforcement learning problems as Markov games and defining a new concept called meta‑NE, the authors establish conditions linking meta‑NE to stationary points of a gradient‑play meta‑MARL algorithm. Experiments on autonomous‑driving tasks show that this approach adapts faster than pretrained MARL baselines, demonstrating its effectiveness.

By Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu
arXiv AI
4d ago

Towards Reliable Vision-Language Models for Autonomous Driving

The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.

By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
arXiv AI
4d ago

Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving

The paper introduces a layered evaluation protocol for generative scenario models used in autonomous driving, focusing on physical consistency and plausibility. It examines internal representations through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis, and then tests outputs against vehicle dynamics constraints such as lateral jerk thresholds. The protocol is applied to a VAE-based scenario generator and other generative models, revealing deeper insights than standard output-level metrics.

By Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
arXiv AI
4d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal
arXiv AI
4d ago

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

STATERA is a method that adapts a pretrained video backbone with mostly frozen weights and a lightweight temporal tubelet mixer to estimate the center-of-mass (CoM) of opaque, asymmetric rigid bodies from short monocular videos. It introduces the HiddenMass Benchmark, consisting of 50K simulated MuJoCo trajectories and a 63-sequence real-world test set with calibrated CoM ground truth. In simulation, STATERA reduces normalized CoM error from 41.7% to 25.2%, and in zero-shot sim-to-real transfer, its phase‑aware variant consistently predicts movement toward the true hidden offset, improving physics capture from 2.6% to 41.0%.

By Animesh Varma
arXiv AI
4d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv Computer Vision
4d ago

Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.

By Changyi Li, Yu Xiao
arXiv Computer Vision
4d ago

CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight

CtrlWAM introduces a controllable world action model that jointly predicts actions (intent) and visual futures (foresight). By executing perturbed actions in a simulator and pairing them with noised visual outcomes, it aligns action predictions with their visual consequences, using warped video–action noise schedules to maintain visual layout responsiveness. The model extends beyond ego‑only control to multiple agent streams, improving action forecasts, video–action agreement, and command following in driving and robotics experiments.

By Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen
arXiv Computer Vision
4d ago

ALFRED: Requirement-driven development of an open-source mobile manipulator for long-term plant monitoring

ALFRED is an open‑source mobile manipulator designed for long‑term plant monitoring, built from commercial parts and featuring a six‑degree‑of‑freedom arm, LiDAR, RGB‑D cameras, RTK GNSS, and an IMU on an Ackermann‑steered base. The platform was iteratively refined over four builds to meet six requirements—durability, modularity, repairability, sensing reach, endurance, and reproducibility—resulting in a 66.1% usable arm reach and clear LiDAR views in the final build. Over a year of monthly forest surveys, ALFRED completed 528 traversals without missing a scheduled collection, demonstrating its reliability despite battery wear and rapid build transitions.

By Ciar\'an Miceal Johnson, Christopher Quail, Garry Ellard, Alistair McConnell, Steve Tonneau, Fernando Auat Cheein
arXiv AI
4d ago

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

The paper introduces TRUST, a token‑level reward model that predicts the correctness of partial chain‑of‑thought (CoT) traces in vision‑language‑action (VLA) policies, enabling monitoring and selective steering of reasoning. On driving and manipulation VLA benchmarks, TRUST improves reasoning accuracy and reduces collision rates and trajectory errors, though its impact on overall task performance varies across tasks. The study defines two evaluation axes—correctability and actionability—to assess when CoT can serve as a runtime safety interface.

By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
arXiv AI
4d ago

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

The paper introduces Divide-and-Remember (D&R), a recursive memory method for vision-language-action (VLA) policies that optimises memory by maximizing the conditional mutual information between actions and memory given observations. D&R recursively divides the full history into top‑K selections over 2K tokens, using a shared lightweight selector across all recursion blocks to handle unbounded histories efficiently. Evaluated on the RoboMME benchmark of 16 long‑horizon manipulation tasks, D&R achieves state‑of‑the‑art success rates with consistent gains across all suites while using only 64 tokens, and similar improvements are observed in real‑robot experiments.

By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
arXiv Machine Learning
4d ago

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

arXiv:2607.04171v4 Announce Type: replace-cross Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...

By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng