arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
The study introduces RegimeShift‑Surrogates, a streaming benchmark that tests surrogate models across eight tasks and multiple regimes. It compares revalidation—choosing the model with lowest current‑window validation loss—to stateful adaptive controllers and finds that revalidation consistently outperforms stateful methods, achieving lower mean log regret in most task‑scenario combinations. The results suggest that fresh validation evidence is more valuable than carrying over past evidence when dealing with distribution shifts.
By Harshil Lodhiya
arXiv:2609.39261v1 Announce Type: new
Abstract: Decision-focused learning (DFL) trains predictors through downstream objectives, but a different loss need not provide an independent parameter-update...
By Aojie Yuan, Haiyue Zhang, Zijian Su
arXiv:2602. 10430v2 Announce Type: replace-cross Abstract: Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures.
By Yusen Huo, Changping Wang, Yangru Huang, Jun Zhang, Jie Jiang
arXiv:2609.39634v1 Announce Type: cross
Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribu...
By Nima H. Siboni