arXiv AI

Fast Generalized Neural Tangent Kernel Statistics via Trace Estimation

The paper presents efficient approximations for key statistics of the Neural Tangent Kernel (NTK) in finite-width neural networks using randomized trace estimation (Hutch++). It demonstrates that the NTK trace, Frobenius norm, effective rank, and alignment can be estimated with high accuracy via matrix-free products, leveraging the NTK’s positive-semidefinite structure to use one-sided estimators with forward or reverse-mode differentiation. Experiments on MLPs, GRUs, and a 410‑million‑parameter Transformer show orders‑of‑magnitude speedups and enable practical state‑space NTK diagnostics at large scales, including applications to RNN training and data‑scarce knowledge distillation.

arXiv Machine Learning
Sep 17

Long-Context Demonstration Selection Using State Space Models

The paper addresses the challenge of selecting demonstrations for long-context language model queries, where transformer inference costs grow quadratically with sequence length. It proposes two algorithms that distill transformer behavior into state space models (SSMs) with linear inference time, partitioning transformer layers into groups and estimating separate SSMs for each. The distilled SSMs achieve less than 0.7% approximation error, and in downstream tasks they reduce FLOPs by 14.2× while improving accuracy by 6.48% compared to baseline methods.

By Ziniu Zhang, Zhenshuo Zhang, Ruoxuan Xiong, Gene Cooperman, Hongyang R. Zhang
arXiv AI
Jul 7

Spectral Signatures of Large Language Models

arXiv:2607. 03377v1 Announce Type: cross Abstract: The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management and quantification at scale, such as model lineage tracing, licensing, and evaluation.

By Zhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu, Hengrui Luo, Pu Ren, Yaoqing Yang
arXiv Machine Learning
Jun 5

Incremental Transformer Neural Processes

arXiv:2602. 18955v2 Announce Type: replace Abstract: Neural Processes (NPs), and specifically Transformer Neural Processes (TNPs), have demonstrated remarkable performance across tasks ranging from spatiotemporal forecasting to tabular data modelling.

By Philip Mortimer, Cristiana Diaconu, Tommy Rochussen, Bruno Mlodozeniec, Richard E. Turner
arXiv AI
Jun 4

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.

By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento