arXiv:2609.36945v1 Announce Type: new
Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version...
By Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding
arXiv:2607. 18830v1 Announce Type: cross Abstract: Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks.
By Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi
arXiv:2409. 03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems.
By El Mahdi Chayti, Martin Jaggi
arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
arXiv:2608. 19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored.
By Zhiqiang Tan
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2609. 30556v1 Announce Type: new Abstract: We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ.
By Naram Mhaisen, George Iosifidis
The paper introduces Network Feasibility Geometry Reinforcement Learning (NFG‑RL), a method that enforces multi‑layer network constraints—such as interference, power‑rate coupling, flow conservation, service chains, capacity, latency, and reliability—by transporting a proto‑policy through a differentiable feasibility map. By compiling heterogeneous constraints into typed residual blocks and using a variational transport operator, NFG‑RL ensures almost‑sure feasible execution and shapes exploration and gradients to respect active constraints. Experiments on two wireless‑edge surrogate environments show that NFG‑RL boosts feasible utility by 37.5–41.5 %, cuts raw‑action violations by 48.5–60.8 %, and reduces P99 delay by 57.0–75.5 % compared to leading baselines.
By Zuyuan Zhang, Zeyu Fang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan
arXiv:2609.07666v1 Announce Type: new
Abstract: Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order...
By Yuyang Wang, Haoyu Yao, Pengcheng Xie
arXiv:2604. 06039v2 Announce Type: replace-cross Abstract: Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL).
By Zhichao Jia, Guanghui Lan
arXiv:2606. 08779v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses.
By Jiashun Liu, Runze Liu, Xu Wan, Jing Liang, Hongyao Tang, Ling Pan
arXiv:2606. 25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints.
By Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal