RIFAR is a new continual learning method for robots that uses reliability screening and drift-aware replay to mitigate forgetting while keeping storage low. It reconstructs past trajectories from short demonstration prefixes and employs a frozen inverse-dynamics model to verify action‑visual consistency. In experiments on LIBERO suites and real‑world tasks, RIFAR outperforms previous generative replay approaches, achieving high performance with only a small fraction of stored steps.
By Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen
UniIntervene++ is an adaptive intervention agent that learns to balance autonomous execution and assisted behaviors during online reinforcement learning. It models the evolving policy, trajectory correction, and task-structured CodePolicy as Options in a unified semi‑Markov decision process, learning their values online. The agent periodically probes the policy to keep control allocation responsive, and uses coupled experience learning to improve both the policy and future interventions, achieving an 89.67% success rate on five real‑world tasks while cutting human intervention by 94.6%.
By Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang
The paper introduces Isaac‑Net, a GPU‑batched 5G New Radio module that integrates a simulated 5G network into massively parallel robot learning environments. It simulates every 0.5 ms slot for thousands of environments simultaneously, matching the median delay and Age of Information of the ns‑3 5G‑LENA simulator while enabling up to one million robots on a single GPU. Extensive experiments show that Isaac‑Net maintains network‑in‑the‑loop performance at 83 % of the Isaac Lab rate without the network, validating its scalability and fidelity.
By Zifan Zhang, Mingzhe Han, Kannan Athreya, Yuchen Liu
CrowdOcc introduces a new RGB‑D dataset and monocular semantic scene completion framework tailored for quadruped robots navigating crowded indoor spaces. The dataset comprises 25.1K frames across 11 scenes, with annotations derived via static‑dynamic decoupling. The framework employs Normal Guided Scene Geometry Fusion and Human‑Centric Sparse Interaction to robustly model occluded geometry and human‑scene relationships, achieving state‑of‑the‑art performance on the dataset’s test set.
By Feiyang Chen, Jincheng Hu, Yiduo Chen, Jihao Li, Yue Liang, Bingzhao Gao, Yanjun Huang, Yuanjian Zhang
The paper examines how multi‑camera streaming perception systems perform on heterogeneous edge platforms that share resources with other workloads. Using two end‑to‑end pipelines on a single GPU–NPU platform, the authors show that isolated single‑stream evaluations can mislead deployment decisions: while the GPU pipeline appears superior in isolation, GPU‑local contention causes deadline misses that make detections stale and can reverse the preferred placement. The study finds that the NPU pipeline, though less accurate for small and medium objects, nearly matches the GPU on large objects, and that under high contention the best placement shifts from All‑GPU to All‑NPU, achieving a 5.2× improvement in worst‑stream sAP. The authors argue that evaluation metrics should include contention sweeps, deadline‑miss rates, and worst‑stream sAP in addition to mean sAP to capture severe single‑stream degradation.
By Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee
SafeStreamingFlow is a goal‑conditioned planner that aligns flow sampling dynamics with execution dynamics by sequentially integrating a learned state vector field with hierarchical state prediction. It enforces safety constraints only for the executed step using high‑order control barrier functions, reducing planning latency and improving safety compared to prior safe diffusion/flow planners. The method demonstrates competitive goal‑reaching success across navigation, racing, and locomotion benchmarks.
By Seunghwan Jang, Jeongyong Yang, Siddharth Ancha, SooJean Han
SCION introduces a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances. By jointly optimizing discrete and continuous scene parameters, SCION achieves high-quality reconstructions at a small size (≈1.2 MB) and outperforms existing Gaussian compression methods in rate‑distortion. The approach enables instance‑level editing and animation without retraining, demonstrating that neural scene representations can discover reusable parts instead of memorizing each element separately.
By William Koch, Amogh Joshi, Cyrus Vachha, Cheng Zheng, Felix Heide
The paper introduces VISTA, a workspace-level equivariant visuotactile diffusion policy designed for data-efficient imitation learning in contact-rich manipulation tasks. VISTA transforms visual and tactile inputs into spherical tokens, fuses them through permutation-equivariant spherical fusion, and aligns the fused representation with the end-effector orientation to produce spatially consistent actions. Experiments in simulation and real-world robotics demonstrate that VISTA significantly improves data efficiency compared to strong visuotactile imitation learning baselines.
By Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou
The paper introduces Bidirectional Voronoi-biased Exploration Curriculum (BVER), a method that expands start states from the goal and goals from the initial state simultaneously, guiding both toward each other to train a single goal-conditioned policy. Inspired by bidirectional RRT planning, BVER biases exploration toward unexplored task space and steers the two expansions together. Experiments on point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer show that BVER learns faster than other reference-free curricula, achieving high success rates and robustness without requiring demonstrations.
By Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter
FUSEye is a lightweight training framework that adapts a frozen COCO‑pretrained YOLO26‑x detector for fisheye images by adding only about 227k parameters. It introduces three modules: GridViews to enlarge compressed boundary regions, Z‑Adapters to correct distortion‑induced feature misalignment, and AgreeFusion to fuse detections across overlapping views. On the WoodScape fisheye benchmark, FUSEye boosts YOLO26‑x mAP50 from 0.148 to 0.266 while requiring only 25% of the labeled data to achieve 0.2597 mAP50, and it also improves YOLOv8‑11 detectors.
By Wenya Su, Kai Luo, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Kunyu Peng, Kailun Yang
The paper introduces a hierarchical model-based reinforcement learning framework that learns control policies from action-free time series by exploiting shared dynamics across related systems. It uses low-dimensional dynamical embeddings to capture both shared structure and individual variation, which then parameterize shared policy and value networks. Experiments on Lorenz-63, double-pendulum, and neural-behavioral data show that these hierarchical policies outperform independently trained ones, achieve higher rewards than planning with the same models, and generalize to unseen systems using only embedding inference.
By Niklas Emonds, Georgia Koppe
The paper introduces a method for quantifying epistemic uncertainty in flow‑matching based generalist robot policies, such as vision‑language‑action models and world‑action models. By measuring velocity‑field disagreement across a small ensemble, the authors obtain better‑calibrated uncertainty estimates that can detect deployment failures and guide active fine‑tuning. Their SAVE approach reduces the need for expert demonstrations, improving real‑world task success from 39 % to 47 % while maintaining a fixed demonstration budget.
By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.
By Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang
The paper introduces an Embodied Neurocomputation framework that tackles the challenge of encoding and decoding between silicon computers and biological neural networks (BNNs). Using a large‑scale parameter sweep, the authors optimized over 1,300 encoding configurations for a BNN agent navigating an odor‑style gradient in a simulated grid‑world, running more than 4,000 hours of real‑time interactions. Twelve configurations consistently outperformed silicon‑based DQN agents within the same interaction budget, demonstrating the framework’s potential for scalable, goal‑oriented learning with BNNs.
By Johnson Zhou, Daniel Tanneberg, Forough Habibollahi, Alon Loeffler, Kiaran Lawson, Valentina Baccetti, Kwaku Dad Abu-Bonsrah, Candice Desouza, Finn Doensen, Bradley Watmuff, Daria Kornienko, Azin Azadi, Justin Leigh Bourke, Bernhard Sendhoff, Brett J. Kagan
WAM-OPD introduces a method for post‑training improvement of World Action Models (WAMs) by collecting student rollout histories and querying a stronger teacher for paired video and action targets. The student learns from both modalities while preserving its one‑step generation capability at deployment. Experiments on 12 RoboTwin 2.0 tasks and four real‑robot tasks show success rates rising from 33.8% to 65.7% and from 51.4% to 64.6%, respectively, with joint video‑action supervision outperforming single‑modality approaches.
By Liuhaichen Yang, Zhengyang Zhong, Hanshang Zhu, Ningwei Bai, Qichen Yin, Zhi Han, Jiarui Qin, Zhuang Jiang, Chenchao Sheng, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Bo Liu, Yi Dong, Zezhi Tang
The paper presents a mechanistic defense for Vision‑Language‑Action (VLA) models against adversarial patches. By using a sparse autoencoder, the authors identify a feature whose activation correlates strongly with the presence of an adversarial patch and suppress this feature only when a linear probe detects an attack. This conditional intervention improves robustness on the LIBERO‑10 benchmark while avoiding the performance degradation that occurs with continuous suppression.
By Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi
arXiv:2610.02527v1 Announce Type: cross
Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong objec...
By Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang
arXiv:2610.02505v1 Announce Type: cross
Abstract: Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient...
By Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil
arXiv:2512.09929v2 Announce Type: replace
Abstract: World models paired with model predictive control (MPC) can be trained offline on large-scale datasets of expert trajectories and enable generaliza...
By Rohun Agrawal, Nimit Kalra, Arjun Parthasarathy, Yann LeCun, Oumayma Bounou, Pavel Izmailov, Micah Goldblum
arXiv:2610.03516v1 Announce Type: cross
Abstract: World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and ac...
By Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li