Learning Length-Extrapolatable Recurrent Models
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates why recurrent models often fail to generalize beyond their training horizon, noting that vanishing or exploding gradients are not the sole cause. It introduces the concept of state credit—the influence of future losses on earlier recurrent states—and proposes Credit Stabilization through Time (CST), a method that rescales this signal during backpropagation to stabilize its norm. Experiments on synthetic and real data show that CST enables models to perform well up to 128 times longer than their training length.
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
Dynamic Compression in Recurrent Networks proposes a method that lets recurrent models revisit and revise their fixed-size state through additional updates, rather than compressing all information in a single causal pass. This approach allows the model to retain lower-fidelity history and refine only the relevant parts when needed, reducing the required state size for accurate task reuse. Experiments show that dynamic compression lowers the recurrent state needed and scales better as the number of stored functions increases.
Dynamic Compression in Recurrent Networks proposes a method for recurrent models to selectively revisit and update past tokens, rather than compressing all history in a single causal pass. By allowing the model to refine its fixed-size state only when needed, it can maintain lower-fidelity information in the raw sequence and revisit it later. Experiments show that this selective re-scanning reduces the recurrent state needed for accurate task reuse and scales better as the number of stored functions increases.
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
arXiv:2608. 07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions.