arXiv AI

Calibration-risk routing for controlled world-model adaptation

The paper introduces the Model-Corrected World Model (MC‑WM), a method for model‑based reinforcement learning that mitigates simulator‑to‑target shift by partitioning initial target data into fit, selection, and calibration sets and choosing the family with the lowest standardized calibration risk. It employs a learned confidence signal and deterministic validity predicates to weight one‑step imagined policy updates, avoiding the need to rewrite physical rewards. The approach is evaluated on 540 unique run cells across three controlled MuJoCo dynamics with contact shifts, with one exact‑routing cell repeated after a pre‑deployment artifact gate, totaling 541 completed executions.

arXiv Machine Learning
Sep 14

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

The paper introduces CLAW, a method that uses a hypernetwork to generate low‑rank adapters for world models during test time, enabling efficient adaptation to new environments with only a few episodes of interaction. By jointly pretraining the hypernetwork and base model on simulated adaptations, CLAW balances computational efficiency and expressivity, outperforming both in‑context learning and gradient‑based adaptation in locomotion and manipulation tasks. The approach also mitigates overfitting in data‑scarce regimes and demonstrates that the benefit stems from expressive adapters rather than context conditioning.

By Fernando Palafox, David Fridovich-Keil
arXiv Machine Learning
Jul 7

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

arXiv:2607. 04978v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning.

By Ruslan Rakhimov, George Bredis, Yuriy Maksyuta, Daniil Gavrilov
arXiv AI
Aug 11

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

arXiv:2608. 09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation.

By Peterson Co, Sicheng Hu, Chunxuan Jiao, Hongyang Cheng, Yulin Luo, Yijie Xu, Sixiang Chen, Zhongxia Zhao, Zihao Wang, DaFeng Chi, Peidong Liu, YuTong Chen, Henghua Liu, Zhihao Yuan, Huizhu Jia, Yuzheng Zhuang, Tianle Zhang, Liang Lin, Huajie Tan, Shanghang Zhang
Hugging Face Trending Papers
Jul 6

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning. A single checkpoint that serves both would defer this choice to inference, when deployment constraints (rollout cost, observation accessibility) determine which path wins.

arXiv Machine Learning
Sep 18

DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

DeliveryGym is a 3D reinforcement learning environment that simulates continuous courier shifts, integrating multimodal tool interaction, persistent world dynamics, and trajectory‑based rewards derived from simulator events. It allows agents to learn how their decisions affect time, energy, and money across an entire shift, and it adapts future training shifts to the policy’s weaknesses while keeping evaluation fixed. Experiments on six models and 13 city maps show a significant gap between task execution and optimal sequencing, with RL improving Qwen3‑VL‑4B’s net income by 54.3% and adaptive training boosting test income by 16.5% over uniform sampling.

By Haoqiang Kang, Yiming Zhang, Yiyang Guo, Chuying Li, Jianzhi Shen, Tianruo Rose Xu, Xiaokang Ye, Lianhui Qin