arXiv:2502.07584v3 Announce Type: replace-cross
Abstract: Many learning algorithms can be represented as Markov processes, and understanding their generalization error is a central topic in learning...
By Benjamin Dupuis, Maxime Haddouche, George Deligiannidis, Umut Simsekli
arXiv:2608.29152v1 Announce Type: cross
Abstract: We study the empirical Sinkhorn estimator of the entropic optimal transport potentials under the uniform loss. Since the potentials are only unique u...
By Denis Belomestny
arXiv:2607. 29675v1 Announce Type: cross Abstract: Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored.
By Arkajyoti Bhattacharjee, Arnab Auddy
arXiv:2606. 19036v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networks.
By Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen, Nhat-Tri Ho, Viet-Hoang Tran, Tho Quan, Tan Minh Nguyen
arXiv:2604. 10727v2 Announce Type: replace-cross Abstract: Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gaussian losses but can be ineffective when losses or rewards are heavy-tailed.
By Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
arXiv:2608.21466v1 Announce Type: new
Abstract: We develop spectral algorithms for selecting state-space partitions that define averaging kernels for finite, ergodic and reversible Markov chains. For...
By Michael C. H. Choi, Youjia Wang
arXiv:2607. 17595v1 Announce Type: new Abstract: We establish mean-square and concentration bounds for stochastic approximation (SA) with arbitrary norm contractive mappings, under a multiplicative noise model where the noise may scale affinely with the norm of the iterates, and the iterates are potentially unbounded.
By Siddharth Chandak
arXiv:2608. 15472v1 Announce Type: cross Abstract: The problem of networked information aggregation, studied in Kearns et al.
By Ambar Pal
arXiv:2606. 02232v1 Announce Type: new Abstract: Learning a Markov transition model is not merely conditional density estimation: the learned object must be a valid transition kernel before it is iterated in downstream dynamics.
By Ao Xu
arXiv:2607. 17232v1 Announce Type: cross Abstract: Classical rate-distortion (RD) theory has long established the fundamental limits of lossy compression by quantifying the minimum number of bits required to represent a source under a prescribed distortion constraint.
By Photios A. Stavrou, Giuseppe Serra, Marios Kountouris
arXiv:2603. 19703v2 Announce Type: replace-cross Abstract: Estimating covariance matrices is fundamental to a wide range of statistical applications.
By T. Tony Cai, Yicheng Li
arXiv:2605. 25085v2 Announce Type: replace-cross Abstract: We study the rate-distortion limits of online KV cache compression in autoregressive language models, formulating it as sequential Wyner-Ziv source coding on the filtration induced by the model, with the next-step query as decoder side information.
By Munsik Kim