arXiv:2608.06481v2 Announce Type: replace-cross
Abstract: Synthesizing safe and robust neural controllers in simulation for reliable sim-to-real deployment remains a critical challenge in robotics. E...
By Riccardo Curcio, Hongpeng Cao, Marco Caccamo
arXiv:2606.20627v2 Announce Type: replace
Abstract: Joint-Embedding Predictive Architectures (JEPAs) enable agents to plan in latent space by imagining the outcomes of candidate actions, yet task spe...
By Samuel Barbeau, Simon Roy, Giovanni Beltrame, Christian Desrosiers, Nicolas Thome
arXiv:2610.02765v1 Announce Type: cross
Abstract: Monitoring planned driving trajectories requires accurately estimating the collision likelihood with actors whose motion is itself impacted by the eg...
By Lu\'is Marques, Rong Fang, Disha Kamale, Dmitry Berenson
arXiv:2610.02666v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert...
By Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee
arXiv:2610.02840v1 Announce Type: cross
Abstract: World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate...
By Chunghyun Park, Beomjun Kim, Seungcheol Park, Heeseung Kwon, Yashu Shukla, Seunghoon Sim, Jinwoo Shin, Minsu Cho
arXiv:2610.03599v1 Announce Type: new
Abstract: High-fidelity 3D head avatars have reached near-photorealistic quality. While recent methods enable text-driven manipulation, they struggle to provide...
By Antonio Canela, Jordi S\`anchez-Riera
arXiv:2610.03248v1 Announce Type: new
Abstract: Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneou...
By Pujun Guo, Yuanfan Zheng, Fei Teng, Mengfei Duan, Guoqiang Zhao, Yuheng Zhang, Kai Luo, Kailun Yang
The paper introduces Forward CIP (F‑CIP), an RL‑native formulation of the Controllable Information Production objective that relies solely on system dynamics, eliminating the need for domain‑specific information variable selection. F‑CIP is proven compatible with reinforcement learning and, when applied to existing algorithms, enables agents to autonomously discover primitive behaviors such as balancing and maintaining controllability. Coupled with a simple forward‑velocity reward, the method yields coordinated gaits like hopping and running without reward engineering.
By Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin
MASkillBlender is a multi-agent reinforcement learning framework that enables decentralized coordination of multiple humanoid robots for loco-manipulation tasks. It learns a shared high-level policy that blends pre-trained single-humanoid skills, requiring only task-level rewards and no task-specific motion references. The approach includes a permutation-based data augmentation technique that preserves policy-gradient direction, and it has been evaluated on several coordination tasks across two humanoid embodiments, consistently achieving strong performance.
By Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang, Chengfeng Jia, Rong Su, Junjie Fu, Guanghui Wen
The paper introduces CRAFT, a method for improving compositional generalization in vision‑language‑action (VLA) models. It addresses the issue where models rely on visual shortcuts during fine‑tuning, leading them to execute demonstrated skill combinations that match observations rather than the instructed ones. By training with counterfactual instruction–observation pairs and transferring supervision through reusable skill representations, CRAFT enhances success on unseen skill combinations while preserving performance on demonstrated ones across multiple VLA models and benchmarks.
By Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon
The paper investigates how reinforcement learning agents trained in one cyber simulation can be transferred to other simulators or real environments. It introduces a framework that decouples state alignment from action translation, allowing zero‑shot policy transfer without retraining. Experiments across four cyber platforms show that transferred policies can preserve performance in closely aligned settings and achieve substantial win rates in more divergent environments.
By Sabrina Saika, Yinuo Du, Aritran Piplai
The paper introduces a continual‑learning framework for single‑view 6‑DoF grasp synthesis with a parallel‑jaw gripper in cluttered scenes. Instead of fine‑tuning a large parametric model, the method updates grasp scores via memory in a learned embedding space and optionally incorporates user demonstrations to generate new candidate grasps. Experiments in simulation and real‑world trials (over 1500 grasps) show that the approach matches baseline performance before adaptation, improves online on unseen objects, and achieves over 90% success on challenging categories after just 50 online attempts.
By Giulio Schiavi, Andrei Cramariuc, Michael Pantic, Roland Siegwart
JEPA-Bisim introduces a bisimulation encoder to joint-embedding predictive world models, ensuring that states with similar transition dynamics are mapped to nearby latent representations while suppressing irrelevant slow features such as background changes and distractors. The approach improves robustness on navigation (PointMaze) and manipulation (PushT) tasks under varied test-time visual conditions, achieving up to tenfold smaller latent spaces than DINO-WM. It remains effective across different pre-trained visual encoders, including DINOv2, SimDINOv2, and iBOT.
By Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Gloria Geng, Nihal Sharma, Donglin Zhan, Nam H. Nguyen, James Anderson
The paper introduces a Frequency-aware Decomposition Network (FDN) that estimates vibration-rich wrench signals in sensorless robotic contact tasks. FDN splits the wrench horizon into low-frequency trends and high-frequency residuals, using pointwise regression for the former and a learned conditional distribution for the latter. Experiments on a 6‑DoF hydraulic manipulator show that FDN reduces high‑frequency amplitude error by up to 47% compared to baselines while maintaining low‑frequency accuracy, and can perform 1,000 ms multi‑step‑ahead estimation in 11 ms on a single CPU thread.
By Hyeonbeen Lee, Min-Jae Jung, Tae-Kyeong Yeu, Jong-Boo Han, Daegil Park, Simon Stepputtis, Jin-Gyun Kim
Variational Streaming Flow (VSF) extends the efficient Streaming Flow (SF) framework by learning a latent distribution conditioned on system dynamics, enabling probabilistic forecasting in physical time. Unlike SF’s deterministic velocity field, VSF produces multiple plausible future trajectories, improving predictive accuracy and distributional fidelity across deterministic and stochastic dynamical systems. The method supports long‑horizon rollouts over 1,000 steps, handles bifurcating dynamics, and can be integrated as a plug‑and‑play predictor into Joint‑Embedding Predictive Architecture (JEPA) world models to enhance navigation, motion planning, and manipulation tasks.
By Hans Hao-Hsun Hsu, Minseon Gwak, Soon Hoe Lim, Pan Li, N. Benjamin Erichson
The paper presents a training-free diffusion-based motion planner that replaces learned global trajectory scores with analytical local scores derived from obstacle, smoothness, velocity, and inter-agent feasibility terms. By reconstructing trajectory scores through local interactions between neighboring waypoints and nearby constraints, the method decomposes the denoising process while preserving the optimization structure of classical trajectory methods. Experiments demonstrate that this approach generates smooth, feasible trajectories for large multi-agent tasks in complex environments quickly, outperforming learning-based and optimization baselines without requiring training data.
By Michael Y. Fatemi, Jinhao Liang, Ferdinando Fioretto
AdaptManip is a fully autonomous framework that enables humanoid robots to navigate, lift, and deliver objects without human demonstrations. It combines a recurrent state estimator for real‑time object tracking, a whole‑body locomotion policy with residual manipulation control, and a LiDAR‑based global position estimator. Trained entirely in simulation with reinforcement learning, the system achieves zero‑shot deployment on real hardware, outperforming imitation‑learning baselines in adaptability and success rate.
By Morgan Byrd, Donghoon Baek, Kartik Garg, Hyunyoung Jung, Daesol Cho, Maks Sorokin, Robert Wright, Sehoon Ha
The paper proposes a reward-based policy that relies only on rewards and actions, enabling zero‑shot transfer between source and target environments with entirely different observation spaces. Experiments on Pointmass, Cartpole, 2D Car Racing, and the Stretch robot in Habitat‑Sim show that the policy can adapt to new visual styles or 3D renderings without additional samples. Additionally, the reward policy can guide the training of an observation‑based policy in the target environment.
By Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha
ScaffoldM3C is a lightweight, multimodal, auto‑regressive framework that generates stable 3D block constructions by treating the task as a probabilistic next‑block generation problem. It incorporates text, image, and sketch conditioning, introduces a scaffold block token to aid intermediate stability, and uses Sequential Monte Carlo to explore multiple assembly sequences simultaneously. The model is four times smaller than existing baselines, achieving 5‑ to 20‑fold inference speedups while matching or surpassing state‑of‑the‑art construction quality and stability in both simulations and real‑world robot demonstrations.
By Gadiel Sznaier Camps, Chengyang He, Guillaume Sartoretti, Eduardo Montijano, Mac Schwager
CF-JEPA is a JEPA-style latent world model that separates the latent space into controllable and uncontrollable subspaces, allowing distractor information to be captured in the uncontrollable region. This factorization prevents latent collapse and maintains performance across 2D and 3D control tasks, even under distracted conditions. The model is validated on a simulated robot task, demonstrating its practical applicability.
By Morgan Byrd, Robert Wright, Sehoon Ha