arXiv Machine Learning

A Thermodynamic Theory of Learning Part II: History-Dependent Reachability and Continual Learning

arXiv Machine Learning
Sep 15

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.

By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
arXiv AI
Aug 7

Continual Learning in Transition

arXiv:2608. 06216v1 Announce Type: cross Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.

By Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
arXiv Machine Learning
Jul 1

Predictable GRPO: A Closed-Form Model of Training Dynamics

arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.

By Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta
arXiv Machine Learning
Sep 14

Transfer Learning for Evolving Domains

The paper introduces Transfer Learning for Evolving Domains (TrED), a framework that models how data availability changes over time in real-world applications. TrED treats the entire trajectory of model updates as a single learning problem, rather than isolated snapshots, and defines a data availability process, a flexible learning protocol, and an evaluation criterion that scores the whole trajectory. The authors review existing transfer learning methods, noting that most are tailored to specific regimes and do not optimize the full trajectory, and argue that TrED is a well‑posed, unsolved research direction.

By Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro, Carlos Soares