arXiv Machine Learning By Mahesh Reddy Pagadala

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

Read the original on arXiv Machine Learning →

The paper reports fine‑grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory‑constrained DNN training. On an LSTM trace, tiny changes in memory budget (0.10% of peak) switch the system between fast and slow execution regimes with up to 7.3× overhead differences, driven by repeated re‑eviction of the same storages. On a ResNet‑32 trace, a deterministic feasibility inversion is observed: the run is feasible at a 0.101 budget ratio, infeasible (OOM) between 0.102–0.106, and feasible again from 0.107, caused by a fully pinned recursive rematerialization frontier exceeding the budget after all evictable tensors are removed. The authors attribute the LSTM instability to the joint size‑staleness scoring term and argue that these represent two distinct budget‑sensitive pathologies rather than a single mechanism.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 25

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

The paper introduces SANE, a method for stabilizing Delta‑Rule recurrent models that maintain a fixed‑size state. By applying adaptive tanh compression at chunk boundaries, SANE prevents localized norm explosions observed in long‑sequence experiments while preserving reasoning performance on short‑context benchmarks. Experiments on a 100M‑token prefix show that SANE retains functional reasoning where the baseline fails, but overly aggressive compression sacrifices reasoning ability, highlighting a capacity–stability trade‑off.

By Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao, Ruichu Cai