arXiv AI By Yifan Zhang, Liang Zheng

Calibration-risk routing for controlled world-model adaptation

Read the original on arXiv AI →

The paper introduces the Model-Corrected World Model (MC‑WM), a method for model‑based reinforcement learning that mitigates simulator‑to‑target shift by partitioning initial target data into fit, selection, and calibration sets and choosing the family with the lowest standardized calibration risk. It employs a learned confidence signal and deterministic validity predicates to weight one‑step imagined policy updates, avoiding the need to rewrite physical rewards. The approach is evaluated on 540 unique run cells across three controlled MuJoCo dynamics with contact shifts, with one exact‑routing cell repeated after a pre‑deployment artifact gate, totaling 541 completed executions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

The paper introduces CLAW, a method that uses a hypernetwork to generate low‑rank adapters for world models during test time, enabling efficient adaptation to new environments with only a few episodes of interaction. By jointly pretraining the hypernetwork and base model on simulated adaptations, CLAW balances computational efficiency and expressivity, outperforming both in‑context learning and gradient‑based adaptation in locomotion and manipulation tasks. The approach also mitigates overfitting in data‑scarce regimes and demonstrates that the benefit stems from expressive adapters rather than context conditioning.

By Fernando Palafox, David Fridovich-Keil
arXiv Machine Learning
Jul 7

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

arXiv:2607. 04978v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning.

By Ruslan Rakhimov, George Bredis, Yuriy Maksyuta, Daniil Gavrilov