arXiv AI By Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao, Ruichu Cai

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

Read the original on arXiv AI →

The paper introduces SANE, a method for stabilizing Delta‑Rule recurrent models that maintain a fixed‑size state. By applying adaptive tanh compression at chunk boundaries, SANE prevents localized norm explosions observed in long‑sequence experiments while preserving reasoning performance on short‑context benchmarks. Experiments on a 100M‑token prefix show that SANE retains functional reasoning where the baseline fails, but overly aggressive compression sacrifices reasoning ability, highlighting a capacity–stability trade‑off.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

The paper introduces DAMP, a decay‑aware mixed‑precision quantization scheme for recurrent‑state representations in GDN and KDA language models. By identifying high‑risk channels through quantization‑error energy and decay persistence, DAMP stores these channels at higher precision while compressing the rest to INT8, achieving a 9.9‑bit average precision. Experiments on Qwen3.6‑35B and Kimi‑Linear‑48B show a 69.1% reduction in recurrent‑state storage, up to 2.01× faster state‑update kernels, and up to 10.9% lower full‑model TPOT while preserving accuracy close to the FP32 baseline.

By Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
arXiv Machine Learning
Aug 20

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

The paper investigates how the dynamical regime of recurrent-depth reasoners—whether they settle, drift, or remain marginal—affects the reliability of test‑time depth. It establishes a depth‑safety condition based on per‑step displacement relative to the decoder margin, showing that operators in a settling regime can safely increase depth without degrading performance and can even improve accuracy on harder unseen tasks such as Sudoku. The authors provide empirical evidence from algorithmic tasks trained on limited data, demonstrate the impact of a terminal fixed‑point objective on depth behavior, and offer operational criteria to identify useful test‑time depth while cataloguing failure modes.

By Ivan Viakhirev, Kirill Borodin, Amirah Almutairi, Serguei Barannikov, Maxim Abramov, Grach Mkrtchian