We study distributionally robust PAC learning for the $0$--$1$-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order $k>1$ and radius $ρ\geq 0$. For hypothesis classes with VC dimension $d$, we establish realizable and agnostic sample-complexity bounds tight up to constant and logarithmic factors, respectively; ordinary empirical risk minimization attains both rates up to logarithmic factors.
arXiv:2608. 13201v1 Announce Type: cross Abstract: We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j).
By Han Dong, Jiaming Li, Yongqiang Gong, Ruixi Li, Yin Liu
arXiv:2607. 16676v1 Announce Type: cross Abstract: How deep does a graph neural network need to be on a sparse graph?
By Aseem Raj Baranwal
arXiv:2606. 07931v1 Announce Type: cross Abstract: We prove a variance-aware pointwise majorizing-measure theorem for centered Gaussian processes.
By Yunbei Xu
arXiv:2606. 07325v1 Announce Type: cross Abstract: We study the minimax rate of estimating a future value $\mu_{t_n+h}$ of a curve $t\mapsto\mu_t$ in the $2$-Wasserstein space $\mathcal{P}_2(\mathbb{R}^d)$ from finitely many noisy snapshots of its past, under an adiabatic bound $\|\nabla_t^k v\|\le\varepsilon$ on the $k$-th covariant derivative of the velocity field.
By Munsik Kim
arXiv:2608. 10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it.
By Ashim Dhor, Pin-Yu Chen
arXiv:2608.10488v2 Announce Type: replace-cross
Abstract: Let $X_1,\ldots,X_n$ be independent Gaussian tensors in $\mathbb{R}^{d_1}\otimes\cdots\otimes\mathbb{R}^{d_k}$ with a common covariance matri...
By Hengzhi He, Guang Cheng
arXiv:2606. 19491v1 Announce Type: new Abstract: Pretrained transformers sit near singular minima of the loss, where the Fisher information metric degenerates along dead directions: directions in parameter space along which the directional Fisher vanishes.
By Tejas Pradeep Shirodkar, P. J. Narayanan
arXiv:2511. 01064v3 Announce Type: replace-cross Abstract: Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions.
By Charles C. Margossian, Isaac E. Rankin, Lawrence K. Saul
arXiv:2609. 01166v1 Announce Type: cross Abstract: The present research is devoted to the nonparametric estimation of a density-dependent drift coefficient in a multivariate McKean--Vlasov diffusion from independent observations at a common time, as well as the stationary density.
By Denis Belomestny, Ekaterina Morozova
arXiv:2608.28150v2 Announce Type: replace
Abstract: How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\e...
By Yuhe Sui, Jianing Zhang, Yingzhi Tang
arXiv:2608. 04686v1 Announce Type: new Abstract: We study distributionally robust PAC learning for the $0$--$1$-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--Read divergence of order $k>1$ and radius $\rho\geq 0$.
By Elad Aigner-Horev, Daniel Rosenberg, Roi Weiss