arXiv:2511.05620v2 Announce Type: replace
Abstract: We study worst-case dynamic regret of specific multi-armed bandit algorithms on piecewise-stationary instances with at most one breakpoint. Our con...
By Gal Mendelson, Eyal Tadmor
arXiv:2609.07162v1 Announce Type: new
Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
By Xin Xu
arXiv:2607. 24662v1 Announce Type: new Abstract: Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap.
By Tianpeng Li, Xuan Guo, Wenjun Wang, Wang Zhang, Pengfei Jiao
arXiv:2608. 09450v1 Announce Type: new Abstract: Betting-based sequential tests and Blackwell approachability are linked by a rate-explicit reduction through support-function residuals.
By Jinze Zhao
Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations.
arXiv:2608. 20337v1 Announce Type: cross Abstract: Accounting for information flow on the path space of trajectories of a nonnegative martingale yields exact variational identities for it, even at arbitrary random times.
By Akshay Balsubramani
arXiv:2609. 01761v1 Announce Type: cross Abstract: A system often has to act long before it learns whether the act worked: a recommender sees a click in seconds and a purchase in days.
By Melika Baghi
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2608.27782v1 Announce Type: cross
Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is t...
By Xujun Che, Depeng Xu, Shuhan Yuan
arXiv:2608. 14020v1 Announce Type: new Abstract: Adding data known to be correct ought to be safe.
By Joseph Sankoorikal Johny
arXiv:2603. 13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget.
By Majid Ghasemi, Mark Crowley
arXiv:2608. 06337v1 Announce Type: cross Abstract: A monotone adversary observes an i.
By Anay Mehrotra