arXiv Machine Learning

When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation

arXiv:2607. 23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map?

arXiv Machine Learning
Aug 20

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

The paper investigates how the dynamical regime of recurrent-depth reasoners—whether they settle, drift, or remain marginal—affects the reliability of test‑time depth. It establishes a depth‑safety condition based on per‑step displacement relative to the decoder margin, showing that operators in a settling regime can safely increase depth without degrading performance and can even improve accuracy on harder unseen tasks such as Sudoku. The authors provide empirical evidence from algorithmic tasks trained on limited data, demonstrate the impact of a terminal fixed‑point objective on depth behavior, and offer operational criteria to identify useful test‑time depth while cataloguing failure modes.

By Ivan Viakhirev, Kirill Borodin, Amirah Almutairi, Serguei Barannikov, Maxim Abramov, Grach Mkrtchian
arXiv Machine Learning
Jun 24

Layer-wise Geometric Approximation Rates for Deep Networks

arXiv:2604. 20219v2 Announce Type: replace Abstract: Depth is widely viewed as a central contributor to the success of deep neural networks, whereas standard neural network approximation theory typically provides guarantees only for the final output and leaves the role of intermediate layers largely unclear.

By Shijun Zhang, Zuowei Shen, Yuesheng Xu
arXiv AI
Sep 18

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

The paper critiques the common practice of evaluating depth usage in depth‑recurrent language models by truncating depth during inference and measuring performance decline. It argues that this method conflates three distinct effects—fewer block applications, reduced computation, and an out‑of‑distribution readout—yet is usually interpreted as measuring only the second. To address this, the authors introduce the Depth Control Protocol (DCP), a suite of positive and negative controls that isolate each factor, along with a training intervention to confirm causality, specifically tailored for depth‑wise weight‑sharing architectures.

By Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung