arXiv:2608.28150v2 Announce Type: replace
Abstract: How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-\(\e...
By Yuhe Sui, Jianing Zhang, Yingzhi Tang
arXiv:2608. 28150v1 Announce Type: new Abstract: Which geometry controls the rank complexity of normalized softmax attention?
By Yuhe Sui, Jianing Zhang
arXiv:2608. 06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work.
By Ibne Farabi Shihab, Joyanta Jyoti Mondal
arXiv:2608. 13922v1 Announce Type: new Abstract: Detecting distributional changes in high dimension is difficult when neither the pre-change nor post-change density is parametrically specified.
By Guoqing Zhang, Zhaixin Chen
We study a tree-structured regularizer over class-prototype layouts in a hierarchical-classification model and ask whether the choice of latent manifold for the prototypes (Euclidean R^d vs. the Poinc...
arXiv:2607. 21761v1 Announce Type: cross Abstract: We prove function-theoretic analogues of a quantitative result of Hodges on extracting the order property from a sufficiently large 2-tree coded in a binary relation.
By G Conant, C Terry
The paper investigates whether placing class prototypes on a hyperbolic manifold (Poincaré ball) rather than a Euclidean space improves the satisfaction of a tree‑structured regularizer in hierarchical classification. Experiments on WikiArt show that hyperbolic prototypes better preserve nearest‑neighbor topology (higher sibling and cousin recall) across multiple tree definitions, while Euclidean prototypes perform similarly to logistic regression on raw features and only hyperbolic models improve local retrieval. The study provides empirical evidence that the choice of latent geometry can affect the fidelity of tree‑structured regularization in real data.
By Peter Flo, Luca Grossmann
arXiv:2605. 14981v2 Announce Type: replace Abstract: Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system.
By Ao Xu, Tieru Wu
arXiv:2609. 20701v1 Announce Type: cross Abstract: We study efficient algorithms for realizing the first-order oracle complexity of optimization of $G$-Lipschitz convex functions with respect to the $\ell_{q}$-norm over an $\ell_{p}$-ball of radius $R$, where $1\leq p,q\leq \infty$.
By David Mart\'inez-Rubio, Crist\'obal Guzm\'an
arXiv:2607. 10194v1 Announce Type: cross Abstract: We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over a compact instruction alphabet $\Sigma_{\mathrm{HG}}$.
By Mario Pascual-Gonzalez, Ezequiel Lopez-Rubio
arXiv:2609.39855v1 Announce Type: new
Abstract: We study the width required for a randomly initialized hidden layer of a neural network to achieve rank lifting. Namely, given a dataset $X \in \mathbb...
By Luca Becchetti, Matteo Russo, Ruben Skorupinski
arXiv:2606. 18520v1 Announce Type: cross Abstract: Computing geometric representations of data is a cornerstone of modern machine learning, typically achieved by training dual encoders which map queries and documents into a shared embedding space.
By Prashant Gokhale, Piotr Indyk, Yuhao Liu, Sandeep Silwal, Tony Chang Wang, Haike Xu