arXiv:2606. 03831v1 Announce Type: new Abstract: This paper investigates non-stationary online learning using the metric of interval regret, which requires an online algorithm to perform well over every time interval.
By Yan-Feng Xie, Shuche Wang, Peng Zhao, Zhi-Hua Zhou
arXiv:2510. 19528v2 Announce Type: replace-cross Abstract: We investigate the fundamental problem of leveraging offline data to accelerate online reinforcement learning - a direction with strong potential but limited theoretical grounding.
By Sebastian Reboul, H\'el\`ene Halconruy
arXiv:2604. 19592v2 Announce Type: replace Abstract: We give a Gordon-Greenwald-Marks (GGM) style black-box reduction from online learning to online multicalibration.
By Gabriele Farina, Juan Carlos Perdomo
arXiv:2605. 14953v2 Announce Type: replace Abstract: We address the problem of conformal selection, where an agent must select a minimal subset of options to ensure that at least one ``success'' is identified with a pre-specified target probability $\phi$.
By Sreenivas Gollapudi, Kostas Kollias, Kamesh Munagala, Ali Sinop
arXiv:2608. 07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications.
By Joar Skalse, Edoardo Pona, Osvaldo Simeone, Nicola Paoletti
arXiv:2510. 15824v2 Announce Type: replace-cross Abstract: This article considers an online version of conformal inference, called adaptive conformal inference [ACI] and introduced by Gibbs and Cand\`es (2021): prediction sets are issued sequentially, after observing features and before the outcomes are revealed.
By Guillaume Principato, Gilles Stoltz
The paper presents an online algorithm that achieves the same $0.401$ approximation factor for maximizing nonnegative, non-monotone DR-submodular functions over compact convex down-closed subsets of the $d$-dimensional unit cube as the best known offline construction. In the full-information value-oracle model, the algorithm attains this factor with sublinear regret, using $O(dT^{1/4})$ oracle calls per round and $O(T^{3/4})$ regret, and offers flexible batching trade-offs. Under a positive-anchor condition, a randomized blocking strategy preserves the $0.401$ factor while achieving $O(T^{5/6})$ one-point bandit regret.
By Vaneet Aggarwal, Yiyang Lu
arXiv:2602. 24207v2 Announce Type: replace Abstract: The use of algorithmic predictions in decision-making leads to a feedback loop where the models we deploy actively influence the data distributions we see, and later use to retrain on.
By Gabriele Farina, Juan Carlos Perdomo
arXiv:2609.05895v1 Announce Type: new
Abstract: We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. In each period, a request ty...
By Menglong Li, Jiawei Zhang
arXiv:2606. 06486v1 Announce Type: new Abstract: In this paper, we study regret minimization in repeated games with \emph{adaptive} opponents who can respond based on histories of play.
By Mingyang Liu, Asuman Ozdaglar, Tiancheng Yu, Kaiqing Zhang
arXiv:2609. 26978v1 Announce Type: cross Abstract: We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set.
By Shinsaku Sakaue
The paper presents an algorithm that lets a learning agent ask for help from a mentor and transfer knowledge between similar states, enabling safe and effective learning in Markov decision processes with irreversible dynamics and infinite state spaces. It proves that both regret and the number of mentor queries grow sublinearly over time, using a sequence of three reductions to achieve a general result. The work claims to be the first formal proof that an agent can achieve high reward while becoming self‑sufficient in an unknown, unbounded, high‑stakes environment without resets.
By Benjamin Plaut, Juan Li\'evano-Karim, Hanlin Zhu, Stuart Russell