Sharp Minimax Regret for Infinite-Memory Logistic Prediction
arXiv:2608. 26515v1 Announce Type: cross Abstract: We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory.
arXiv:2606. 21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally.
arXiv:2608. 26515v1 Announce Type: cross Abstract: We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory.
arXiv:2608. 07922v1 Announce Type: new Abstract: Adaptive learning needs both a state that preserves what observations imply and opportunities to act on that state.
arXiv:2608.27782v1 Announce Type: cross Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is t...
arXiv:2609. 30556v1 Announce Type: new Abstract: We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ.
arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version...
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
arXiv:2512. 23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$.
arXiv:2608. 10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using.
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2606. 04182v1 Announce Type: cross Abstract: We formulate the problem of \emph{exact unlearning} in reinforcement learning, where the goal is to design an efficient framework that enables the removal of any user's data upon deletion request, i.
arXiv:2609.37660v1 Announce Type: new Abstract: We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each roun...
arXiv:2608. 06337v1 Announce Type: cross Abstract: A monotone adversary observes an i.