arXiv Machine Learning

Trajectory Design and Budgeted Querying for Digital Twin Calibration

arXiv:2608. 08631v1 Announce Type: new Abstract: Digital-twin calibration requires interaction data that is expensive to collect.

arXiv AI
Sep 10

Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

The paper introduces an action‑conditioned world‑modeling framework that turns Earth‑system simulator trajectories into training data for controllable state‑transition learning. By pretraining on naturally observed state changes as implicit action supervision and using masked response learning, the model can infer unobserved variables and learn coupled system dependencies. Experiments on ecosystem dynamics across six global regions demonstrate that the model maintains long‑horizon emulation accuracy while enabling structural interventions and coherent responses in coupled ecosystem‑cycle variables.

By Zhihao Wang, Ruichen Wang, Ruohan Li, Lei Ma, George Hurtt, Xiaowei Jia, Gengchen Mai, Shaowen Wang, Yiqun Xie
arXiv Machine Learning
Sep 25

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

The paper proposes a method for selectively querying language‑model advice in reinforcement learning by predicting the value of potential responses and only querying when the expected benefit outweighs the cost. It introduces a certified, response‑contingent metareasoning framework that guarantees near‑optimal advice usage under certain assumptions, and demonstrates that a calibrated controller with Qwen2.5 advisors can improve task performance while drastically reducing the number of advice calls on the BabyAI benchmark.

By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv Machine Learning
Aug 17

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.

By Anisa Saleem, Duksu Kim
arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal