arXiv:2606. 03003v1 Announce Type: cross Abstract: A latent world model built from an equivariant encoder $E$ and an equivariant predictor $f$ inherits a provable symmetry of its training loss: when the world's dynamics genuinely carries a group $G$ acting on latents by an orthogonal representation $\rho(g)$, the one-step prediction relMSE is exactly invariant across the whole group, so fitting the dynamics on a restricted slice of orientations mathematically determines it on the entire orbit (j\v{u} y\=i f\v{a}n s\=an).
By Hongbo Wang (Stony Brook University)
arXiv:2608.28150v2 Announce Type: replace
Abstract: How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\e...
By Yuhe Sui, Jianing Zhang, Yingzhi Tang
arXiv:2607. 20578v1 Announce Type: new Abstract: We study Gaussian-width complexity on statistical manifolds through a pair of functionals: the primal Fisher width $w_G(T) = w(G^{1/2}T)$, induced by the Fisher metric, and the inverse-Fisher width $w_{G^{-1}}(T) = w(G^{-1/2}T)$, induced by the inverse Fisher metric.
By Vu Khac Ky
arXiv:2608.29702v1 Announce Type: new
Abstract: A token-embedding table holds a hub of short rows near its origin, and we show that this cluster biases what nearest-neighbor intrinsic-dimension (ID)...
By Alexandre Quemy
arXiv:2609.22337v1 Announce Type: cross
Abstract: Linear probing is the standard instrument for detecting social biases in the hidden representations of large language models. Yet reported probe accu...
By Mo Hai, Haifeng Li
arXiv:2609.39855v1 Announce Type: new
Abstract: We study the width required for a randomly initialized hidden layer of a neural network to achieve rank lifting. Namely, given a dataset $X \in \mathbb...
By Luca Becchetti, Matteo Russo, Ruben Skorupinski
The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
By Christopher M. Bryant, Hao Liu
arXiv:2608. 08103v1 Announce Type: new Abstract: Smooth acyclicity constraints answer whether a weighted support is a DAG, whereas structure learning asks which support change should be made.
By Rui Wu, Zongyuan Chen, Hong Xie
arXiv:2607. 15702v1 Announce Type: cross Abstract: We prove a finite-sample formulation gap for physics-informed learning of nonlinear multiscale elliptic equations.
By Ronald Katende
arXiv:2608. 25326v1 Announce Type: new Abstract: In transductive classification, an adversary fixes a labeled population, one label is hidden uniformly, and the learner sees all remaining labels.
By Pahan Dewasurendra
arXiv:2606. 01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse.
By Triet M. Le
arXiv:2604. 14727v2 Announce Type: replace Abstract: To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention.
By Ye Su, Yong Liu