arXiv Machine Learning

Statistical Gains from Looped Estimation under Parameter Budgets

The paper investigates whether a looped estimator—one that repeatedly applies a single fitted operator with shared parameters—can enhance statistical accuracy while staying within a fixed parameter budget. It establishes upper and lower bounds on squared Hellinger risk for looped sieve maximum likelihood and compares them to the untied counterpart, revealing a tradeoff between parameter sharing, iteration count, and accuracy. For models with known H"older smoothness, looped residual feedforward networks and a post‑layer‑normalized Transformer achieve minimax polynomial rates with a fixed number of bounded real parameters, and under certain conditions the looped estimator’s worst‑case risk vanishes as sample size grows, outperforming the untied approach.

arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv Machine Learning
Aug 11

Scalable extensions to given-data Sobol' index estimators

arXiv:2509. 09078v3 Announce Type: replace-cross Abstract: Given-data methods for variance-based sensitivity analysis have significantly advanced the feasibility of Sobol' index computation for computationally expensive models and models with many inputs.

By Teresa Portone, Bert Debusschere, Samantha Yang, Emiliano Islas-Quinones, T. Patrick Xiao
arXiv AI
2d ago

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.

By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu