Robotics and embodied AI

Manipulation, locomotion, sim-to-real transfer and autonomous driving: learning systems that have to survive physics.

3,855 stories · RSS feed

arXiv AI
Sep 25

Rolling-WAM: World Action Models with Rolling Imagination

Rolling-WAM is a new formulation for World Action Models that spreads the joint video-action denoising process across multiple replanning cycles. It keeps a sliding window of video-action chunks at different noise levels, fully denoising the immediate chunk for execution while partially refining future chunks. This approach reduces latency, improves closed-loop responsiveness, and achieves a 4.5× speedup in steady-state replanning compared to standard WAMs while maintaining competitive manipulation performance.

By Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
arXiv AI
Sep 25

RAPID: Robot Agentic Programming from Demonstrations

RAPID is a system that automatically generates, verifies, and refines robot programs from a single visual human demonstration. It infers a testable task specification, action primitives, and an interactive environment, using an object-centric relational program representation to enable reuse beyond the demonstration. The approach was evaluated in simulation on eight contact-rich manipulation tasks and successfully deployed on a real Franka arm, showing strong generalization across object pose, shape, material, and environment.

By Yuyao Liu, Jiayuan Mao, David Hsu, Leslie Pack Kaelbling, Tom\'as Lozano-P\'erez
arXiv AI
Sep 25

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

CrossSafe proposes embodiment-conditioned safety filtering that uses a Hamilton‑Jacobi reachability value function shared across robots while conditioning on each robot’s morphology and kinematics via a morphology‑aware latent representation. The method performs reachability analysis directly in latent space, enabling a single policy trained on multiple bimanual robot embodiments and manipulation tasks to generalize zero‑shot to a held‑out embodiment and reduce collision rates. Experiments on five embodiments and five tasks demonstrate that training with more embodiments improves generalization.

By Ihab Tabbara, Yuxuan Yang, Hussein Sibai
arXiv AI
Sep 25

Band-Attention Modulation Network for Robust Face Forgery Detection

The paper introduces Band-Attention Modulation Network (BAM‑Net), a face forgery detection framework that learns fine‑grained, adaptive modulation of frequency bands in the Discrete Cosine Transform spectrogram. BAM‑Net dynamically reweights anti‑diagonal frequency bands to enhance forgery‑related spectral cues while suppressing irrelevant information, then fuses this modulated frequency data with spatial features using a lightweight backbone with distance‑decayed attention. Experiments on FaceForensics++, Celeb‑DF, and DFDC show that BAM‑Net achieves state‑of‑the‑art performance and strong generalization across datasets, compression levels, and manipulation types.

By Zhida Zhang, Wenkui Yang, Xinlei Ma, Qihang Fan, Jie Cao
arXiv AI
Sep 25

RoboLDA: A Probabilistic Generative Model for Uncovering Embodied Hierarchical Structures in Voxel-based Soft Robots

RoboLDA is a Bayesian probabilistic model that learns a four‑level hierarchy—task, robot, organ, voxel—from existing high‑performing voxel‑based soft robot designs. By training with variational inference, it uncovers consistent, intuitive hierarchical patterns and can generate new robot morphologies that outperform evolutionary algorithms by an average of 106.4% without further optimization. The inferred organ structures also improve modular control policies, demonstrating the model’s utility for zero‑shot design and motion control.

By Junru Song, Yang Yang, Jingdan Shi, Guozhen Li, Weien Zhou, Ying Wen, Feifei Wang, Wen Yao, Tingsong Jiang
arXiv AI
Sep 25

S2Planner: Multi-Scale Semantic Planner for End-to-End Autonomous Driving

S2Planner is a trajectory planner for autonomous driving that fuses data from three front-facing cameras, ego‑motion history, and the current driving command. It uses a fine‑tuned DINOv3 backbone with a Spatial Tuning Adapter to generate multi‑scale image features, which are refined by a coarse‑to‑fine decoder employing trajectory self‑attention and camera‑projected cross‑attention. The key contribution lies in integrating ego‑conditioned trajectory initialization with iterative, geometry‑guided sampling of multi‑scale image features, rather than introducing a new visual backbone or attention operator.

By Zhaowei Lu, Liguo Zhou, Yujie Guo, Lei Yu, Alois Knoll
arXiv AI
Sep 25

MorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots

MorphIK is a flow‑matching neural model that learns inverse kinematics for revolute‑joint kinematic chains it has never seen during training. Using a transformer to encode a robot’s morphology and target pose, the model generates pose solutions from noise and can be fine‑tuned with optimization to achieve sub‑centimeter accuracy. It also efficiently samples the robot’s null space, producing diverse configurations for the same pose.

By Lennart Clasmeier, Jan Gerrit Habekost, Cornelius Weber, Stefan Wermter
arXiv AI
Sep 25

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

World Action Agent (WAA) is a multi‑agent framework that lets vision‑language models (VLMs) directly pilot robots by operating within a visual action workspace. The workspace provides automatically selected contact views, editable action rehearsals, and in‑view correction to refine decisions before low‑level execution. WAA learns procedural skills from expert videos and human teaching, and its interaction traces can train smaller VLMs, achieving state‑of‑the‑art success on LIBERO‑Pro and improving out‑of‑domain performance on robosuite and Qwen3.5‑9B.

By Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li
arXiv Computer Vision
Sep 25

Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

The paper introduces FlipDir, a training‑free inference‑time technique that mitigates answer flips in vision‑language models by steering hidden states along a low‑rank subspace derived from contrastive image pairs. It employs a margin‑based gate to attenuate steering only during uncertain decoding steps, thereby restoring original predictions while keeping stable ones unchanged. The authors also present VisFlip, a benchmark framework that evaluates models across nine dataset‑variation combinations in scientific reasoning, robot‑scene understanding, and medical VQA, showing that FlipDir consistently outperforms existing methods on recovery and preservation metrics.

By Yeonsung Jung, Joonhyun Jeong, Hoang Pham, Joowon Kim, Yoonsik Park, Viet Dac Lai, Eunho Yang
arXiv Machine Learning
Sep 25

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

The paper introduces Predictive Action Chunk Learning (PACL), a method for improving robot manipulation policies using mixed-quality deployment experience. PACL first trains a predictive chunk-level critic to evaluate temporally extended action sequences, then uses the critic’s quality estimates to guide a diffusion actor that learns from both successful and failed rollouts. Experiments on simulated and real robots demonstrate that PACL consistently enhances pretrained policies and outperforms strong imitation learning and offline reinforcement learning baselines.

By Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, Chen Lv
arXiv AI
Sep 25

BEHAVE: Real-Time Modeling of Human Systems as Observable Complex Dynamical Systems and Operational Objects for Physical AI

BEHAVE models an interacting human group as a complex dynamical system called a HumanSystem, whose state is partly encoded in the interaction structure rather than individual tracks. By incorporating interaction evidence, the method improves group discrimination and captures differences in neighbor-level organization during bottleneck scenarios. The framework derives routing, local dynamics, and stability metrics, enabling real-time querying of group state and critical modes for Physical AI applications.

By Helene Malyutina
arXiv Machine Learning
Sep 25

Safety-oriented pedestrian trajectory prediction at urban intersections using time-to-collision and crossing-zone context

The paper introduces a safety‑oriented pedestrian trajectory prediction framework for urban intersections that fuses pedestrian motion history with Time‑to‑Collision (TTC) data and crossing‑zone context. Using the inD dataset, a pooled LSTM architecture encodes TTC and context separately before integrating them with pedestrian positions, and a weighted loss emphasizes large errors. The approach reduces average displacement error (ADE) and final displacement error (FDE) and significantly lowers the frequency of errors exceeding a 1 m tolerance compared to a position‑only model.

By Erel Avineri, Yftach Gil, Yehudit Aperstein
arXiv Machine Learning
Sep 25

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

The paper investigates task collapse—a failure mode where online RL fine‑tuning of a pretrained flow‑matching vision‑language‑action policy erodes performance on individual tasks—using a 450M‑parameter SmolVLA policy on LIBERO‑10. Three exploration‑noise strategies are compared: a fixed noise scale, a learned noise network, and an uncertainty‑gated controller that reallocates exploration based on novelty and competence signals without task labels. The uncertainty‑gated controller prevents task collapse across all tested seeds, whereas the other two approaches consistently cause collapse, demonstrating its effectiveness in preserving task performance during fine‑tuning.

By Mehmet Turan Yard{\i}mc{\i}, Yunus Emre \c{C}o\u{g}urcu
arXiv Machine Learning
Sep 25

Continuous Online Fault Detection for Mobile Robots via Adaptive Edge Models

The paper presents a Teacher-Student distillation framework for continuous online fault detection in mobile robots. An offline foundation model (TSPulse) generates pseudo‑labels from augmented time‑series data, while a lightweight MiniRocket Student, enhanced with a Recursive Least Squares estimator, performs real‑time inference with a 4.30 ms CPU latency. The Student adapts online to domain shifts, improving VUS‑PR scores from 0.26 to 0.75 and uses an uncertainty‑guided active learning strategy to request minimal operator interventions.

By Jordan Levy, Nicolas Verstaevel, Vincent Talon, Benoit Gaudou
arXiv Computer Vision
Sep 25

DeltaWAM: Delta World Action Models for Bimanual Manipulation

DeltaWAM introduces a new approach to world-action models (WAMs) for bimanual manipulation by jointly predicting visual deltas and actions instead of dense future frames, thereby reducing redundant modeling of unchanged content and mitigating nuisance appearance variations. The method employs three architectures with varying representation and computation sharing, and incorporates Streaming Delta Memory (SDM) to update cached anchor context using compact observed deltas, which cuts heavy video-expert processing. Experiments on RoboTwin show that DeltaWAM with SDM raises average success rates from 81.3% to 85.4% in clean settings and from 75.8% to 83.9% under visual randomization, while also reducing training FLOPs by up to 23.77% and inference latency by 36.57%. whyItMatters":"DeltaWAM improves both performance and computational efficiency for bimanual manipulation tasks by focusing on visual deltas and efficient memory updates, as demonstrated by higher success rates and lower FLOPs on RoboTwin."

By Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang
arXiv Computer Vision
Sep 25

IronViT: Toward Efficient Generalist Visual Representation Learning

IronViT proposes a new approach to building efficient generalist vision encoders by first consolidating the knowledge of multiple specialist teachers into a softmax attention bridge and then transferring this consolidated representation to a hybrid softmax‑linear attention architecture. This two‑stage distillation process, supported by a curated data pipeline, allows the model to capture semantic, spatial, language‑aligned, and action‑relevant cues while avoiding the high‑resolution cost of traditional softmax attention. Across tasks such as recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT matches or exceeds the performance of leading specialist and generalist encoders, with the hybrid encoder offering increasing efficiency at higher resolutions.

By Jiaxi Huang, Yueqi Hu, Xin Zhu, Xiaopeng Zhang, Huiting Qiao, Yanglin Zhang, Zefeng Ji, Rongxue Li, Yifei Xu, Huiying Yu, Wei Liu, Jiayin Zheng, Yinggan Xu, Peipeng Chen, Yin Zhang, Jian Yao
arXiv Computer Vision
Sep 25

Representation World Model: Learning States, Transition and Executable Plans in Representation

The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.

By Yijun Yuan, Weicheng Zheng, Weibang Wang, Minghui Qin, Chang Sun, Junhao Huang, Kenan Li, Anmin Liu, Yicheng Yao, Hang Zhao
arXiv Computer Vision
Sep 25

Free-Init: Scan-Free, Motion-Free, and Correspondence-Free Initialization for Doppler LiDAR-Inertial Systems

The paper introduces Free-Init, a high‑frequency, resilient initialization framework for LiDAR‑inertial systems that uses FMCW Doppler LiDAR to capture both point range and Doppler velocity. By fusing point‑wise Doppler velocity with inertial measurements, Free‑Init eliminates the need for motion undistortion, excitation motions, and map correspondences during initialization, making it plug‑and‑play for a wide range of initial motions, including stationary, dynamic, and violent movements. Experiments on diverse platforms and motion scenarios demonstrate that Free‑Init achieves fast convergence and high‑frequency performance, delivering outputs exceeding 10 kHz and outperforming existing methods.

By Mingle Zhao, Jiahao Wang, Tianxiao Gao, Chengzhong Xu, Hui Kong
arXiv Computer Vision
Sep 25

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

BeyondRetarget is an end‑to‑end framework that learns to generate executable humanoid robot motions directly from monocular RGB videos, bypassing the need for an explicit human motion representation. By learning robot‑oriented implicit representations and incorporating a contact‑aware motion optimization mechanism, the method captures cross‑morphology motion structures and improves temporal consistency and physical plausibility. Experiments demonstrate that BeyondRetarget achieves higher execution success rates, lower latency, and greater accuracy and robustness in both simulation and real humanoid robots.

By Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao
arXiv Computer Vision
Sep 25

Self-Adaptive VLA for Robust Robot Deployment

The paper introduces Self‑Adaptive VLA, a post‑training method that lets Vision‑Language‑Action policies self‑adapt to deployment‑time hardware shifts by using rollouts as context. It creates shift‑conditioned expert demonstrations, compresses visual, proprioceptive, and action data into a latent context token, and modulates the policy via adaptive layer normalization. Experiments on four precision‑critical manipulation tasks show the method recovers over 80 % of the base policy’s performance under actuation bias and encoder offsets, and improves robustness on new workstations.

By Hongxin Zhang, Chunru Lin, Tsun-Hsuan Wang, Zhenjia Xu, Chuang Gan