arXiv:2609.23775v1 Announce Type: new
Abstract: Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we...
By Suman Banerjee, Hiroyasu Tsukamoto
The paper introduces an adaptive rollout truncation method for offline world model training that uses epistemic uncertainty to decide when to stop autoregressive rollouts. By calibrating a threshold during a warm‑up phase, the approach replaces fixed‑horizon rollouts with uncertainty‑driven truncation, evaluated with ensemble and Monte Carlo dropout estimators. Experiments on ANYmal‑D and ANT demonstrate that this strategy matches or surpasses fixed‑horizon training while reducing cumulative rollout steps by about 72%.
By Nikodem Sebastian Zymla, Laurin Thiele, Johannes Pitz
arXiv:2608. 02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress.
By Inkyu Sa, Konstantin Stulov, Rajat Bhageria
arXiv:2605.22164v2 Announce Type: replace
Abstract: Latent world models can learn representations that contain information needed for control, while the downstream controller may still rank candidate...
By Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu
arXiv:2607. 27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators.
By Takumi Shioda, Kohei Terashima, Tatsuo Nagai
arXiv:2503. 03660v4 Announce Type: replace Abstract: We introduce a sequence-conditioned critic for Soft Actor-Critic (SAC) that models trajectory context with a lightweight Transformer and trains on aggregated $N$-step targets.
By Dong Tian, Onur Celik, Gerhard Neumann