VLA-Scope is a two‑stage framework designed to predict failures in vision‑language‑action models under distribution shifts. The first stage detects out‑of‑distribution inputs and classifies their shift categories using pooled image and language representations. For OOD inputs, the second stage updates failure risk during execution by combining shift category, action‑prefix features, and execution progress, achieving a ROC‑AUC of 0.8497 after 60 actions and outperforming baseline methods.
By Kaiwen Zhu, Dongfang Liu, Liangkai Liu
AtomEgo investigates how to integrate large-scale egocentric human interaction data into embodied foundation model pre‑training. The study uses a curated 2,659‑hour corpus and a scalable data pipeline to evaluate three co‑training paradigms across vision‑language‑action and world‑action architectures. Results show that the benefit of egocentric data depends on both its scale and the quality of alignment with robotic embodiment, offering practical guidance for scalable ego‑robot pre‑training.
By Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie, Xiaoquan Sun, Junyang Zheng, Zhuoyang Song, Jiaxing Zhang, Jiayu Chen
The paper introduces PA‑RL, a reinforcement‑learning framework that uses artificial potential fields as the action representation for contact‑rich robotic manipulation. Instead of directly commanding motion, the policy adjusts potential‑field parameters, which a Cartesian impedance controller then executes, decoupling task strategy from low‑level control. In peg‑in‑hole experiments, PA‑RL achieved a 100% success rate in simulation, outperformed baselines in torque and acceleration variation, and transferred to a real robot without fine‑tuning.
By Xinyu Liu, G\"okhan Solak, Arash Ajoudani
SynthDemo‑RL introduces a teacher‑student framework that uses an automated teacher to generate successful manipulation trajectories from simulator‑privileged state, which are then distilled into a Vision‑Language‑Action (VLA) student via supervised fine‑tuning. The student is further refined with PPO using binary task‑success rewards. On the LIBERO‑PRO benchmark, SynthDemo‑RL rescues all 27 previously unsolvable tasks and achieves near‑perfect success rates, matching performance that would otherwise require human demonstrations.
By Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo, Kenji Fukumizu, Manohar Kaul
The paper examines whether fine‑tuning large language models (LLMs) with personality‑labelled data improves their ability to act as socially interactive agents. Two small open‑weight LLMs were fine‑tuned on a corpus of personality‑labelled social media posts and dialogues, and the resulting models were evaluated in various social interaction scenarios by independent LLM judges. The findings show that the fine‑tuned models do not outperform their baseline counterparts in role‑playing personalities, though they offer comparable text quality and increased linguistic diversity for the Qwen models; low inter‑rater agreement limits confidence in the results, suggesting future work should focus on training data quality and domain alignment.
By Tim Krabbe, Xiaodan Shi
arXiv:2609.21942v1 Announce Type: cross
Abstract: A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt...
By Eshika Pathak, Leela Krishna
MemeLens is a unified multilingual, multitask Vision‑Language Model designed to improve meme understanding across a wide range of tasks such as hate, misogyny, propaganda, sentiment, and humour. The authors consolidated 38 public meme datasets, mapping their labels into a shared taxonomy of 20 tasks covering harm, targets, figurative intent, and affect, and conducted extensive experiments to show that multimodal training and a unified approach outperform fine‑tuning on individual datasets. All experimental resources, the model, and the datasets are released publicly for community use.
By Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam
The paper presents a comparative study of motion planning methods from major autonomous driving benchmarks—CARLA, nuPlan, and the Waymo Open Dataset—using CARLA Leaderboard v2.1 as a unified evaluation platform. Eight representative planners (TF++, InterFuser, TCP, PDM‑Lite, MTR+MPC, CaRL, PlanT 2.0, Diffusion planner) are evaluated to highlight their strengths, weaknesses, prevailing trends, and common challenges in motion planning research.
By Merve Atasever, Alfredo Reina Corona, Zhuochen Liu, Qingpei Li, Akshay Hitendra Shah, Hans Walker, Jyotirmoy V. Deshmukh, Rahul Jain
arXiv:2604.07799v3 Announce Type: replace-cross
Abstract: Robots deployed for long periods keep improving their skills, and each update changes a released system. We treat this as a software-lifecycl...
By Xue Qin, Simin Luan, Cong Yang, Zhijun Li
arXiv:2606.03963v4 Announce Type: replace-cross
Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but still relies heavily on time consuming manual re...
By Roohan Ahmed Khan, Yasheerah Yaqoot, Amir Atef Habel, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou
The paper introduces a training‑free proactive defense for detecting partial deepfake speech by using self‑embedding steganography. It embeds a compressed version of the clean audio within itself, allowing post‑hoc extraction of reference content and enabling detection of spoofed segments via codec‑based restoration. Experiments on a benchmark dataset show that this method complements passive detectors and operates without any training, offering a robust, data‑efficient alternative for partial deepfake detection.
By Yigitcan \"Ozer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi
FootQuery is a perceptive locomotion framework that retrieves depth information from a robot’s own history by querying each foot’s predicted next touchdown. The policy uses proprioceptive predictions of touchdown locations and uncertainties to sample relevant historical depth frames, fuses these per‑foot features with global visual memory, and generates control actions. In simulation and on a real Unitree G1 robot, FootQuery enables continuous traversal of complex outdoor stairs, indoor routes, platforms, and gaps, outperforming component ablations.
By Tao Dong, Jia Yu, Yuxuan Fan, Linna Zhao, Jiaqi Gong, Andong Yang, Chao Gao, Guyue Zhou
The study examines how different vision‑language‑action (VLA) policies execute a manipulation task by comparing the geometry of their end‑effectors across 15,000 closed‑loop LIBERO rollouts. By pairing 3,600 configuration‑matched policy executions, the authors find that when both policies succeed, their end‑effector trajectories are much closer (median DTW distance 0.0120 m) than when only one succeeds (0.0380 m), a pattern consistent across all tasks, policy pairs, and nine representations. Even successful executions remain as far from same‑task demonstrations as the demonstrations are from each other, indicating that task‑associated geometry, rather than training data overlap, drives these differences.
By Xingyu Lin, Zhuang Li, Zhongrun Wu, Shouquan Zhou, Dehui Du
MT‑WAM enhances the Fast‑WAM framework by adding complementary supervision for future 2‑D point trajectories and visual features while keeping the original training objectives. A lightweight dual‑stream branch and structured attention mask isolate motion‑specific processing, and motion‑stream tokens provide additional dynamics cues to the action expert. During inference, MT‑WAM skips future‑video prediction, using cached video and motion information to achieve higher success rates on LIBERO, LIBERO‑Plus, RoboTwin 2.0 Clean2Rand, and several real‑world tasks.
By Yiguang Yang, Jiankun Peng, Xiaoming Wang, Yiran Zhang, Zhibo Fang
The paper introduces a compositional continual learning benchmark for world models in robot manipulation, designed to isolate knowledge reuse from learning speed and capacity. Tasks are curated to combine previously seen action and perception components, allowing analysis of how different modalities affect reuse. Experiments show that modular world models better balance reuse and forgetting than conventional methods, yet none fully solve the challenge, highlighting the need for models explicitly built to reuse knowledge without forgetting.
By Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner
The paper introduces an adaptive rollout truncation method for offline world model training that uses epistemic uncertainty to decide when to stop autoregressive rollouts. By calibrating a threshold during a warm‑up phase, the approach replaces fixed‑horizon rollouts with uncertainty‑driven truncation, evaluated with ensemble and Monte Carlo dropout estimators. Experiments on ANYmal‑D and ANT demonstrate that this strategy matches or surpasses fixed‑horizon training while reducing cumulative rollout steps by about 72%.
By Nikodem Sebastian Zymla, Laurin Thiele, Johannes Pitz
The paper introduces PARTS, a real‑world subtask reinforcement learning framework that fine‑tunes a pretrained robot policy by focusing on critical bottleneck subtasks while keeping the base policy frozen. It uses agent‑generated selectors and success verifiers to provide local rewards, enabling learning even when full‑task successes are rare. Experiments on bimanual YAM and single‑arm Franka robots show that PARTS raises complete‑task success from 32% to 61% and from 50% to 95%, respectively, with only tens of minutes of real‑world RL rollouts and minimal human intervention.
By Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun
HERMES is a holistic end‑to‑end multimodal driving framework that incorporates long‑tail semantic knowledge into trajectory planning for autonomous vehicles. It uses a foundation‑model‑assisted annotation pipeline to build Long‑Tail Scene Context and Long‑Tail Planning Context, capturing hazard‑centric scene information, maneuver intent, and risk‑aware guidance. A Tri‑Modal Driving Module then fuses multi‑view visual observations, historical ego‑motion, and long‑tail semantic instructions to generate intent‑ and risk‑aware trajectories, achieving consistent performance gains on a large‑scale real‑world long‑tail driving benchmark.
By Weizhe Tang, Junwei You, Jiaxi Liu, Zhaoyi Wang, Rui Gan, Zilin Huang, Feng Wei, Bin Ran
arXiv:2609.21221v1 Announce Type: new
Abstract: Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical ru...
By Hongyan Wei, Wael AbdAlmageed
arXiv:2609.21751v1 Announce Type: cross
Abstract: Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these...
By Tim Engelbracht, Ren\'e Zurbr\"ugg, Mayank Mittal, Marco Hutter, Marc Pollefeys, Hermann Blum, Zuria Bauer