arXiv Machine Learning

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B. whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."

arXiv Machine Learning
Sep 10

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

The paper introduces GaugeLasso, a method that applies symmetric group‑lasso penalties to transformer channels during training, enabling entire tensor slices to be zeroed out while maintaining dense tensors for GPU efficiency. By calibrating channel penalties based on inference utility per compute, the network self‑organizes into depth‑dependent structural profiles that can be dramatically smaller than the original architecture, achieving up to 255‑fold compression on a polynomial division task and outperforming hand‑designed baselines on language modeling and autoencoding benchmarks. The approach also accelerates training and reveals over‑provisioned axes that guide subsequent design iterations.

By Jed A. Duersch, Na\"im Es-Sebbani, Nathana\"el Haas, Zied Bouraoui
arXiv Machine Learning
Sep 11

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

LILA (Latent-Informed Layer Analysis) introduces a calibration‑free method for structured pruning of large language models by scoring neuron importance using the Kolmogorov–Smirnov distance between singular value distributions of full and neuron‑ablated feed‑forward network weight matrices. The approach requires no training, calibration data, or auxiliary networks, and outperforms existing methods such as PruneNet and SliceGPT on LLaMA‑2‑7B and Phi‑2 at various sparsity levels. After a single epoch of LoRA fine‑tuning, LILA matches heavily calibrated baselines, and a Neural Tangent Kernel analysis provides theoretical support for its spectral importance criterion. Additionally, LILA can dynamically allocate sparsity budgets, achieving state‑of‑the‑art generative preservation and revealing architectural bottlenecks at higher compression.

By Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad
arXiv Machine Learning
6d ago

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

Spectral-Guided Diffusion introduces a method to accelerate diffusion inference by identifying and reusing residual branches that need not be recomputed during the trajectory. The approach uses a Spectral Concentration Ratio (SCR) combined with Frobenius magnitude to create an offline sensitivity proxy and deterministic lifetime for each scheduled unit, eliminating the need for routers or input-dependent searches. Experiments on models such as LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL show that this scheduling preserves quality better than several baselines and achieves up to a 3.0× wall‑clock speedup over eager inference.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma
arXiv Computation and Language
6d ago

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

The study re‑examines a reported advantage of a routed ternary (1.58‑bit) language model over a full‑precision transformer at 60K parameters. By running controlled experiments with multiple seeds and a fixed training recipe, the authors find that the apparent benefit largely stems from the choice of baseline model shape rather than the ternary architecture itself. While the routed model does outperform other shapes at a larger 130M‑byte budget, its advantage diminishes when a plain gated diagonal‑SSM block is used, and the ternary penalty varies with architecture and quantization details.

By Gautam Veldanda
arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra