arXiv:2606. 03073v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models (LLMs) is highly sensitive to hyperparameter configurations, making hyperparameter optimization (HPO) essential yet computationally expensive.
By Minping Chen, Bowen Xiao, Du Liang, Chuxuan Zeng, Zeyi Wen
arXiv:2607. 26680v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks.
By Mingxuan Che, Tsung-Yuan Tseng, Theresa Eimer, Marius Lindauer, Alexander von Rohr
arXiv:2606. 02521v1 Announce Type: new Abstract: One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization.
By Zhou Jiang, Yandong Wen, Zhen Liu
The paper introduces Repeated Optimizer Resampling (ROR), a method that treats optimizer choice as a hyperparameter and searches for the best optimizer during a single training run. ROR periodically scouts each candidate optimizer for a short number of epochs, then continues training with the best scout, allowing the optimizer to change over time. Experiments on MNIST, Fashion‑MNIST, and motor insurance claim‑count models show that one‑epoch ROR uses only 24–35% of the training effort required to exhaustively evaluate all optimizers while achieving comparable performance.
By Ronald Richman, Mario V. W\"uthrich
The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.
By Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa
One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization. We propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for deterministic one-step generators.
arXiv:2509. 23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent.
By Matt L. Sampson, Peter Melchior
Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations.
arXiv:2504. 13822v3 Announce Type: replace-cross Abstract: The emergence of large pre-trained networks has revolutionized the AI field, unlocking new possibilities and achieving unprecedented performance.
By Eric Nuertey Coleman, Luigi Quarantiello, Ziyue Liu, Qinwen Yang, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco
arXiv:2609.36659v1 Announce Type: new
Abstract: The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these st...
By Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang, Wei Lin, Guojun Yin, Qingming Huang, Shuhui Wang
The paper investigates how reusing past samples can improve the sample efficiency of Proximal Policy Optimization (PPO). Two variants, wPPO-U and wPPO-BH, are introduced within a multiple importance weighting framework, each reusing data from recent iterations while preserving core PPO mechanics. The authors derive theoretical policy improvement bounds for both variants and empirically evaluate their impact on continuous control tasks.
By Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli
arXiv:2601. 19810v2 Announce Type: replace-cross Abstract: Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks.
By Octavio Pappalardo