arXiv AI

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

The paper introduces SANE, a method for stabilizing Delta‑Rule recurrent models that maintain a fixed‑size state. By applying adaptive tanh compression at chunk boundaries, SANE prevents localized norm explosions observed in long‑sequence experiments while preserving reasoning performance on short‑context benchmarks. Experiments on a 100M‑token prefix show that SANE retains functional reasoning where the baseline fails, but overly aggressive compression sacrifices reasoning ability, highlighting a capacity–stability trade‑off.

arXiv Machine Learning
Aug 31

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

The paper introduces DAMP, a decay‑aware mixed‑precision quantization scheme for recurrent‑state representations in GDN and KDA language models. By identifying high‑risk channels through quantization‑error energy and decay persistence, DAMP stores these channels at higher precision while compressing the rest to INT8, achieving a 9.9‑bit average precision. Experiments on Qwen3.6‑35B and Kimi‑Linear‑48B show a 69.1% reduction in recurrent‑state storage, up to 2.01× faster state‑update kernels, and up to 10.9% lower full‑model TPOT while preserving accuracy close to the FP32 baseline.

By Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
arXiv Machine Learning
Aug 20

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

The paper investigates how the dynamical regime of recurrent-depth reasoners—whether they settle, drift, or remain marginal—affects the reliability of test‑time depth. It establishes a depth‑safety condition based on per‑step displacement relative to the decoder margin, showing that operators in a settling regime can safely increase depth without degrading performance and can even improve accuracy on harder unseen tasks such as Sudoku. The authors provide empirical evidence from algorithmic tasks trained on limited data, demonstrate the impact of a terminal fixed‑point objective on depth behavior, and offer operational criteria to identify useful test‑time depth while cataloguing failure modes.

By Ivan Viakhirev, Kirill Borodin, Amirah Almutairi, Serguei Barannikov, Maxim Abramov, Grach Mkrtchian
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
arXiv Machine Learning
5d ago

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

The paper reports fine‑grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory‑constrained DNN training. On an LSTM trace, tiny changes in memory budget (0.10% of peak) switch the system between fast and slow execution regimes with up to 7.3× overhead differences, driven by repeated re‑eviction of the same storages. On a ResNet‑32 trace, a deterministic feasibility inversion is observed: the run is feasible at a 0.101 budget ratio, infeasible (OOM) between 0.102–0.106, and feasible again from 0.107, caused by a fully pinned recursive rematerialization frontier exceeding the budget after all evictable tensors are removed. The authors attribute the LSTM instability to the joint size‑staleness scoring term and argue that these represent two distinct budget‑sensitive pathologies rather than a single mechanism.

By Mahesh Reddy Pagadala