Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
22h ago

RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning

RIFAR is a new continual learning method for robots that uses reliability screening and drift-aware replay to mitigate forgetting while keeping storage low. It reconstructs past trajectories from short demonstration prefixes and employs a frozen inverse-dynamics model to verify action‑visual consistency. In experiments on LIBERO suites and real‑world tasks, RIFAR outperforms previous generative replay approaches, achieving high performance with only a small fraction of stored steps.

By Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen
arXiv Machine Learning
22h ago

UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning

UniIntervene++ is an adaptive intervention agent that learns to balance autonomous execution and assisted behaviors during online reinforcement learning. It models the evolving policy, trajectory correction, and task-structured CodePolicy as Options in a unified semi‑Markov decision process, learning their values online. The agent periodically probes the policy to keep control allocation responsive, and uses coupled experience learning to improve both the policy and future interventions, achieving an 89.67% success rate on five real‑world tasks while cutting human intervention by 94.6%.

By Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang
arXiv AI
22h ago

Network-in-the-Loop at Scale: GPU-Batched 5G Simulation for Massively Parallel Robot Learning

The paper introduces Isaac‑Net, a GPU‑batched 5G New Radio module that integrates a simulated 5G network into massively parallel robot learning environments. It simulates every 0.5 ms slot for thousands of environments simultaneously, matching the median delay and Age of Information of the ns‑3 5G‑LENA simulator while enabling up to one million robots on a single GPU. Extensive experiments show that Isaac‑Net maintains network‑in‑the‑loop performance at 83 % of the Isaac Lab rate without the network, validating its scalability and fidelity.

By Zifan Zhang, Mingzhe Han, Kannan Athreya, Yuchen Liu
arXiv Computer Vision
22h ago

CrowdOcc: Monocular Semantic Scene Completion for Quadruped Robots in Crowded Indoor Environments

CrowdOcc introduces a new RGB‑D dataset and monocular semantic scene completion framework tailored for quadruped robots navigating crowded indoor spaces. The dataset comprises 25.1K frames across 11 scenes, with annotations derived via static‑dynamic decoupling. The framework employs Normal Guided Scene Geometry Fusion and Human‑Centric Sparse Interaction to robustly model occluded geometry and human‑scene relationships, achieving state‑of‑the‑art performance on the dataset’s test set.

By Feiyang Chen, Jincheng Hu, Yiduo Chen, Jihao Li, Yue Liang, Bingzhao Gao, Yanjun Huang, Yuanjian Zhang
arXiv Computer Vision
22h ago

Lightweight and Resource-Efficient Perception for Robotic Guide Dogs

The paper examines how multi‑camera streaming perception systems perform on heterogeneous edge platforms that share resources with other workloads. Using two end‑to‑end pipelines on a single GPU–NPU platform, the authors show that isolated single‑stream evaluations can mislead deployment decisions: while the GPU pipeline appears superior in isolation, GPU‑local contention causes deadline misses that make detections stale and can reverse the preferred placement. The study finds that the NPU pipeline, though less accurate for small and medium objects, nearly matches the GPU on large objects, and that under high contention the best placement shifts from All‑GPU to All‑NPU, achieving a 5.2× improvement in worst‑stream sAP. The authors argue that evaluation metrics should include contention sweeps, deadline‑miss rates, and worst‑stream sAP in addition to mean sAP to capture severe single‑stream degradation.

By Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee
arXiv Machine Learning
22h ago

Safe Streaming Flow Planning by Aligning Sampling Dynamics with Execution Dynamics

SafeStreamingFlow is a goal‑conditioned planner that aligns flow sampling dynamics with execution dynamics by sequentially integrating a learned state vector field with hierarchical state prediction. It enforces safety constraints only for the executed step using high‑order control barrier functions, reducing planning latency and improving safety compared to prior safe diffusion/flow planners. The method demonstrates competitive goal‑reaching success across navigation, racing, and locomotion benchmarks.

By Seunghwan Jang, Jeongyong Yang, Siddharth Ancha, SooJean Han
arXiv Computer Vision
22h ago

SCION: Scene Composition with Instanced Neural Primitives

SCION introduces a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances. By jointly optimizing discrete and continuous scene parameters, SCION achieves high-quality reconstructions at a small size (≈1.2 MB) and outperforms existing Gaussian compression methods in rate‑distortion. The approach enables instance‑level editing and animation without retraining, demonstrating that neural scene representations can discover reusable parts instead of memorizing each element separately.

By William Koch, Amogh Joshi, Cyrus Vachha, Cheng Zheng, Felix Heide
arXiv AI
22h ago

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

The paper introduces VISTA, a workspace-level equivariant visuotactile diffusion policy designed for data-efficient imitation learning in contact-rich manipulation tasks. VISTA transforms visual and tactile inputs into spherical tokens, fuses them through permutation-equivariant spherical fusion, and aligns the fused representation with the end-effector orientation to produce spatially consistent actions. Experiments in simulation and real-world robotics demonstrate that VISTA significantly improves data efficiency compared to strong visuotactile imitation learning baselines.

By Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou
arXiv Machine Learning
22h ago

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

The paper introduces Bidirectional Voronoi-biased Exploration Curriculum (BVER), a method that expands start states from the goal and goals from the initial state simultaneously, guiding both toward each other to train a single goal-conditioned policy. Inspired by bidirectional RRT planning, BVER biases exploration toward unexplored task space and steers the two expansions together. Experiments on point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer show that BVER learns faster than other reference-free curricula, achieving high success rates and robustness without requiring demonstrations.

By Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter
arXiv Computer Vision
22h ago

FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

FUSEye is a lightweight training framework that adapts a frozen COCO‑pretrained YOLO26‑x detector for fisheye images by adding only about 227k parameters. It introduces three modules: GridViews to enlarge compressed boundary regions, Z‑Adapters to correct distortion‑induced feature misalignment, and AgreeFusion to fuse detections across overlapping views. On the WoodScape fisheye benchmark, FUSEye boosts YOLO26‑x mAP50 from 0.148 to 0.266 while requiring only 25% of the labeled data to achieve 0.2597 mAP50, and it also improves YOLOv8‑11 detectors.

By Wenya Su, Kai Luo, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Kunyu Peng, Kailun Yang
arXiv Machine Learning
22h ago

Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings

The paper introduces a hierarchical model-based reinforcement learning framework that learns control policies from action-free time series by exploiting shared dynamics across related systems. It uses low-dimensional dynamical embeddings to capture both shared structure and individual variation, which then parameterize shared policy and value networks. Experiments on Lorenz-63, double-pendulum, and neural-behavioral data show that these hierarchical policies outperform independently trained ones, achieve higher rewards than planning with the same models, and generalize to unseen systems using only embedding inference.

By Niklas Emonds, Georgia Koppe
arXiv Machine Learning
22h ago

Uncertainty Quantification for Flow-Based Generalist Robot Policies

The paper introduces a method for quantifying epistemic uncertainty in flow‑matching based generalist robot policies, such as vision‑language‑action models and world‑action models. By measuring velocity‑field disagreement across a small ensemble, the authors obtain better‑calibrated uncertainty estimates that can detect deployment failures and guide active fine‑tuning. Their SAVE approach reduces the need for expert demonstrations, improving real‑world task success from 39 % to 47 % while maintaining a fixed demonstration budget.

By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
arXiv Computer Vision
22h ago

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

The paper introduces IG‑VLA, a vision‑language‑action framework that learns to imagine task‑relevant future scene evolution in a latent spatiotemporal space, guiding action prediction without generating full pixel‑level videos. It further compresses this future reasoning into a compact Scene Gist Token via Scene Gist Memory, enabling efficient inference. Experiments on LIBERO, LIBERO‑Plus, and VLABench show that IG‑VLA improves success rates by nearly 6% and speeds up inference up to 6.38× compared to strong baselines.

By Shenglan Li, Zhendong Mi, Hengyi Zhu, Jingwu Luo, Chun Kit Chan, Geng Yuan, Yanzhi Wang, Pu Zhao, Shaoyi Huang
arXiv Machine Learning
22h ago

Embodied Neurocomputation: A Framework for Interfacing Biological Neural Cultures with Scaled Task-Driven Validation

The paper introduces an Embodied Neurocomputation framework that tackles the challenge of encoding and decoding between silicon computers and biological neural networks (BNNs). Using a large‑scale parameter sweep, the authors optimized over 1,300 encoding configurations for a BNN agent navigating an odor‑style gradient in a simulated grid‑world, running more than 4,000 hours of real‑time interactions. Twelve configurations consistently outperformed silicon‑based DQN agents within the same interaction budget, demonstrating the framework’s potential for scalable, goal‑oriented learning with BNNs.

By Johnson Zhou, Daniel Tanneberg, Forough Habibollahi, Alon Loeffler, Kiaran Lawson, Valentina Baccetti, Kwaku Dad Abu-Bonsrah, Candice Desouza, Finn Doensen, Bradley Watmuff, Daria Kornienko, Azin Azadi, Justin Leigh Bourke, Bernhard Sendhoff, Brett J. Kagan
arXiv AI
22h ago

WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation

WAM-OPD introduces a method for post‑training improvement of World Action Models (WAMs) by collecting student rollout histories and querying a stronger teacher for paired video and action targets. The student learns from both modalities while preserving its one‑step generation capability at deployment. Experiments on 12 RoboTwin 2.0 tasks and four real‑robot tasks show success rates rising from 33.8% to 65.7% and from 51.4% to 64.6%, respectively, with joint video‑action supervision outperforming single‑modality approaches.

By Liuhaichen Yang, Zhengyang Zhong, Hanshang Zhu, Ningwei Bai, Qichen Yin, Zhi Han, Jiarui Qin, Zhuang Jiang, Chenchao Sheng, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Bo Liu, Yi Dong, Zezhi Tang
arXiv AI
22h ago

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

The paper presents a mechanistic defense for Vision‑Language‑Action (VLA) models against adversarial patches. By using a sparse autoencoder, the authors identify a feature whose activation correlates strongly with the presence of an adversarial patch and suppress this feature only when a linear probe detects an attack. This conditional intervention improves robustness on the LIBERO‑10 benchmark while avoiding the performance degradation that occurs with continuous suppression.

By Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi