arXiv AI By Zhiyuan Chen, Yuxuan Zhong, Fan Wang, Bo Yu, Pengtao Shao, Shaoshan Liu, Ning Ding

StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation

Read the original on arXiv AI →

arXiv:2603. 23571v2 Announce Type: replace-cross Abstract: Effective navigation intelligence relies on long-term memory to support both immediate generalization and sustained adaptation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

Extending LLM Context via Associative Recurrent Memory

arXiv:2607. 11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling.

By Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov
arXiv Computer Vision
Sep 25

Visual Representation and History Modeling for Navigation World Models

The paper introduces Navigation World Models (NWMs) that predict action‑conditioned visual futures for planning. It examines two key design challenges: choosing an effective visual representation and efficiently modeling observation history for repeated queries. The authors propose a conditional flow‑transformer framework, compare five frozen visual representations, and develop Cached‑Linear and Balanced GDN architectures to reduce redundant history computation and improve memory usage, demonstrating their effectiveness on RECON, SACSoN, and SCAND benchmarks.

By Guangfu Guo, Xiaoqian Lu, Rui Liu, Yutong Chen, Kunpeng Liu, Long Cheng
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen