The paper introduces GAVA, a grounded arbitration framework that enables text-based embodied agents to decide whether to accept, reject, inspect, or ask for clarification when a user’s correction might be incorrect. In the ALFWorld environment, GAVA achieves perfect correction accuracy through local inspections and reduces interaction costs compared to always-verify baselines, especially when leveraging semantic priors. The study demonstrates that selective information gathering can lower declared joint costs, though it does not conclusively prove a general advantage of environmental value of information over clarification.
By Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin
The paper introduces a meta-multi-agent reinforcement learning (meta‑MARL) framework that enables rapid adaptation of interactive policies in multi‑agent systems. By modeling multi‑agent reinforcement learning problems as Markov games and defining a new concept called meta‑NE, the authors establish conditions linking meta‑NE to stationary points of a gradient‑play meta‑MARL algorithm. Experiments on autonomous‑driving tasks show that this approach adapts faster than pretrained MARL baselines, demonstrating its effectiveness.
By Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu
The paper evaluates five vision‑language models on autonomous driving tasks under various visual input conditions, finding that visual corruption affects accuracy and confidence differently across models and datasets. It then tests Visual Evidence Augmentation (VEA) as an inference‑time technique to enhance reliability, observing mixed improvements depending on the model and setting.
By Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
The paper introduces a layered evaluation protocol for generative scenario models used in autonomous driving, focusing on physical consistency and plausibility. It examines internal representations through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis, and then tests outputs against vehicle dynamics constraints such as lateral jerk thresholds. The protocol is applied to a VAE-based scenario generator and other generative models, revealing deeper insights than standard output-level metrics.
By Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.
By Aryan Goyal
STATERA is a method that adapts a pretrained video backbone with mostly frozen weights and a lightweight temporal tubelet mixer to estimate the center-of-mass (CoM) of opaque, asymmetric rigid bodies from short monocular videos. It introduces the HiddenMass Benchmark, consisting of 50K simulated MuJoCo trajectories and a 63-sequence real-world test set with calibrated CoM ground truth. In simulation, STATERA reduces normalized CoM error from 41.7% to 25.2%, and in zero-shot sim-to-real transfer, its phase‑aware variant consistently predicts movement toward the true hidden offset, improving physics capture from 2.6% to 41.0%.
By Animesh Varma
DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.
By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.
By Changyi Li, Yu Xiao
CtrlWAM introduces a controllable world action model that jointly predicts actions (intent) and visual futures (foresight). By executing perturbed actions in a simulator and pairing them with noised visual outcomes, it aligns action predictions with their visual consequences, using warped video–action noise schedules to maintain visual layout responsiveness. The model extends beyond ego‑only control to multiple agent streams, improving action forecasts, video–action agreement, and command following in driving and robotics experiments.
By Chensheng Peng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, Yuxiao Chen
ALFRED is an open‑source mobile manipulator designed for long‑term plant monitoring, built from commercial parts and featuring a six‑degree‑of‑freedom arm, LiDAR, RGB‑D cameras, RTK GNSS, and an IMU on an Ackermann‑steered base. The platform was iteratively refined over four builds to meet six requirements—durability, modularity, repairability, sensing reach, endurance, and reproducibility—resulting in a 66.1% usable arm reach and clear LiDAR views in the final build. Over a year of monthly forest surveys, ALFRED completed 528 traversals without missing a scheduled collection, demonstrating its reliability despite battery wear and rapid build transitions.
By Ciar\'an Miceal Johnson, Christopher Quail, Garry Ellard, Alistair McConnell, Steve Tonneau, Fernando Auat Cheein
The paper introduces TRUST, a token‑level reward model that predicts the correctness of partial chain‑of‑thought (CoT) traces in vision‑language‑action (VLA) policies, enabling monitoring and selective steering of reasoning. On driving and manipulation VLA benchmarks, TRUST improves reasoning accuracy and reduces collision rates and trajectory errors, though its impact on overall task performance varies across tasks. The study defines two evaluation axes—correctability and actionability—to assess when CoT can serve as a runtime safety interface.
By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
The paper introduces Divide-and-Remember (D&R), a recursive memory method for vision-language-action (VLA) policies that optimises memory by maximizing the conditional mutual information between actions and memory given observations. D&R recursively divides the full history into top‑K selections over 2K tokens, using a shared lightweight selector across all recursion blocks to handle unbounded histories efficiently. Evaluated on the RoboMME benchmark of 16 long‑horizon manipulation tasks, D&R achieves state‑of‑the‑art success rates with consistent gains across all suites while using only 64 tokens, and similar improvements are observed in real‑robot experiments.
By Xuehui Yu, Eason Yu, Meiyi Wang, Haozhe Du, Stefano V. Albrecht, Harold Soh
arXiv:2602.02269v2 Announce Type: replace-cross
Abstract: We present $multipanda\_ros2$, a novel open-source ROS2 architecture for multi-robot control of Franka Robotics robots. Leveraging ros2 contr...
By Jon \v{S}kerlj, Seongjin Bien, Abdeldjallil Naceri, Sami Haddadin
arXiv:2603.16065v3 Announce Type: replace-cross
Abstract: Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecke...
By Yanru Wu, Weiduo Yuan, Esteban Martinez Licon, Ang Qi, Vitor Guizilini, Jiageng Mao, Yue Wang
arXiv:2609.38764v1 Announce Type: new
Abstract: Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract,...
By Zachary Shinnick, Hemanth Saratchandran, Damien Teney, Anton van den Hengel
arXiv:2609.40149v1 Announce Type: new
Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can dif...
By Seonvin Cho, Soohyun Choi, Songnam Hong
arXiv:2503.10118v3 Announce Type: replace-cross
Abstract: The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world syst...
By Yuxuan Xu, Shiyu Wang, Jinhao Huang, Wenhao Zhao, Yufei Jia, Zike Yan, Weibin Gu, Lu Shi, Guyue Zhou
arXiv:2602.23934v2 Announce Type: replace-cross
Abstract: This paper presents a novel autonomous robotic assembly framework for constructing stable structures without relying on predefined architectu...
By Jingwen Wang, Johannes Kirschner, Paul Rolland, Luis Salamanca, Stefana Parascho
arXiv:2607.04171v4 Announce Type: replace-cross
Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...
By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv:2610.01210v1 Announce Type: new
Abstract: Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, wh...
By Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao