arXiv:2606. 27032v1 Announce Type: cross Abstract: Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints.
By Jesper Klicks, Sander Vr\v{z}ina, Vincent Fran\c{c}ois-Lavet
The paper introduces Feedback‑Enriched Environments (FEEs) as a new approach to training large language models as autonomous agents for long‑horizon tasks. By shifting from action guidance to observation enrichment during later stages of exploration, FEEs improve performance across SciWorld and BFCL benchmarks with various Qwen3 model scales and RL algorithms. The study shows that FEEs stabilize training, promote proactive exploration, embed environmental guidance into policy weights, and highlight intra‑group feedback consistency as key for stable optimization.
By Hongbang Yuan, Zhuoran Jin, Yixin Cao
arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.
By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
The study couples the Met Office Unified Model with distributed reinforcement learning agents, using a DDPG actor that applies bounded potential‑temperature corrections across 70 vertical levels. Training is performed on ten nudged forecasts, after which the frozen policy is evaluated in a non‑nudged forecast, demonstrating numerical stability. The learned policy reduces Z₅₀₀ MAE in four of six latitude bands—up to 45.8% in the northern tropics—and decreases MSLP error by up to 27.3% in certain bands, indicating promising bias‑correction potential.
By Pritthijit Nath, Sebastian Schemm, Peter Haynes, Emily Shuckburgh, Mark Webb
HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.
By Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
arXiv:2609.36864v1 Announce Type: new
Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling com...
By Fanchao Chen, Hengyu Fu, Shivaram Venkataraman, Jiantao Jiao
arXiv:2607. 21644v1 Announce Type: new Abstract: We present a goal-agnostic control framework for partial differential equations (PDEs) built around a joint-embedding predictive architecture (JEPA).
By Jonathan Gallagher, Roberto Guglielmi
arXiv:2607. 14180v1 Announce Type: cross Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset.
By Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.
By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv:2608.21946v1 Announce Type: cross
Abstract: Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exp...
By Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
arXiv:2607. 16204v1 Announce Type: new Abstract: Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments.
By Darshan Deshpande