arXiv AI

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?

arXiv AI
Jul 22

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

arXiv:2607. 18323v1 Announce Type: cross Abstract: Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by recomputation after each mutation.

By Abdallah Khemais (ISITCOM, University of Sousse)
arXiv Machine Learning
Sep 22

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B. whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."

By Chenyu Zhu, Ruoyu Zhao, Zhichao Lu
arXiv Machine Learning
Sep 23

Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World

The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.

By Christopher M. Bryant, Hao Liu
arXiv Machine Learning
Sep 10

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge

The paper presents a block‑wise differentiable Sinkhorn attention mechanism designed for long‑context balanced entropic optimal transport on TPU hardware. By stopping a $T$‑step Sinkhorn solve and unrolling a short refinement tail, the authors derive an exact surrogate gradient that achieves efficient block‑wise cost and memory usage. Experimental results on synthetic masked problems and a Pfam protein‑family screen demonstrate high numerical accuracy and sustained throughput on TPU v6e‑8, with notable improvements in reconstruction and sparse cross‑entropy metrics.

By Dylan Forde
arXiv Machine Learning
Sep 25

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

Spectral-Guided Diffusion introduces a method to accelerate diffusion inference by identifying and reusing residual branches that need not be recomputed during the trajectory. The approach uses a Spectral Concentration Ratio (SCR) combined with Frobenius magnitude to create an offline sensitivity proxy and deterministic lifetime for each scheduled unit, eliminating the need for routers or input-dependent searches. Experiments on models such as LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL show that this scheduling preserves quality better than several baselines and achieves up to a 3.0× wall‑clock speedup over eager inference.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma
arXiv Machine Learning
Aug 4

TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention

arXiv:2608. 02050v1 Announce Type: cross Abstract: Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the model's behaviour?

By Avni Mittal, Avinash Anand, Ashutosh Kumar, Dikshant Kukreja, Kritarth Prasad, Sushane Dulloo, Erik Cambria, Timothy Liu, Zhengkui Wang, Rajiv Ratn Shah