The paper introduces the Progressive Memory Transformer (PMT), a transformer variant that adds writable, window‑aligned memory to expose mid‑range representations alongside token and sequence‑level outputs. PMT is trained with a hierarchical learning framework that applies separate objectives at local, mid‑range, and global scales, encouraging the model to capture fine‑grained variation, window‑level motifs, and overall sequence agreement. Experiments on seven UCR/UEA/UCI classification datasets, a cue‑retention probe, and forecasting tasks show that PMT achieves strong low‑label classification performance, competitive multi‑horizon forecasting, and evidence that its memory states encode mid‑range motifs.
By Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
By Shota Takashiro, Masanori Koyama, Takeru Miyato, Yusuke Iwasawa, Yutaka Matsuo, Kohei Hayashi
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
By Akarsh Kumar, Phillip Isola
arXiv:2609.39082v1 Announce Type: new
Abstract: As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and...
By Wentao Wang, Hengyu Zhong, Yunhan Jiang, Jialiang An, Meng Lu
arXiv:2511. 20577v5 Announce Type: replace Abstract: Real-world time series often exhibit strong non-stationarity, complex nonlinear dynamics, and behavior expressed across multiple temporal scales, from rapid local fluctuations to slow-evolving long-range trends.
By Sumit S Shevtekar, Chandresh K Maurya
The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
By Ke Wan, Chen Chen
arXiv:2610.01192v1 Announce Type: new
Abstract: Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between...
By Yi Chen, MingMing Yu, Rui-Qi Wang, Boran Wang, Xiaohang Cao, Chu Tang, Jingmin Chen, Jie Gu
arXiv:2602. 18131v2 Announce Type: replace Abstract: Temporal Predictive Coding provides a layer-local, parallelisable mechanism for learning in recurrent systems, making it an attractive candidate for online local learning on neuromorphic and edge hardware.
By Tom Potter, Oliver Rhodes
arXiv:2607. 23284v1 Announce Type: new Abstract: Automated sleep staging is increasingly used in large-scale studies to derive sleep-architecture endpoints: total sleep time, REM latency, sleep efficiency, and bout-duration statistics.
By Juntang Wang, Yihan Wang, Hao Wu, Jiayu Gao, Shixin Xu, Dongmian Zou
arXiv:2609.37587v1 Announce Type: cross
Abstract: Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulatio...
By Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu, Zuozhu Liu
arXiv:2609.38356v1 Announce Type: new
Abstract: Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continu...
By Sima Hashemi, Daniel Durstewitz, Georgia Koppe
arXiv:2602. 07628v2 Announce Type: replace Abstract: While the shift toward unified foundation models has revolutionized many deep learning domains, sleep medicine remains largely restricted to task-specific models that focus on localized micro-structure features.
By Keondo Park, Younghoon Na, Yourim Choi, Hyunwoo Ryu, Hyun-Woo Shin, Hyung-Sin Kim