arXiv AI

StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation

arXiv:2603. 23571v2 Announce Type: replace-cross Abstract: Effective navigation intelligence relies on long-term memory to support both immediate generalization and sustained adaptation.

arXiv AI
Jul 14

Extending LLM Context via Associative Recurrent Memory

arXiv:2607. 11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling.

By Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov
arXiv Computer Vision
Sep 25

Visual Representation and History Modeling for Navigation World Models

The paper introduces Navigation World Models (NWMs) that predict action‑conditioned visual futures for planning. It examines two key design challenges: choosing an effective visual representation and efficiently modeling observation history for repeated queries. The authors propose a conditional flow‑transformer framework, compare five frozen visual representations, and develop Cached‑Linear and Balanced GDN architectures to reduce redundant history computation and improve memory usage, demonstrating their effectiveness on RECON, SACSoN, and SCAND benchmarks.

By Guangfu Guo, Xiaoqian Lu, Rui Liu, Yutong Chen, Kunpeng Liu, Long Cheng
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen
arXiv AI
Sep 3

AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers

The paper introduces the AGI Maze Prediction Datasets and Benchmark, a lightweight, procedurally generated grid‑world testbed for evaluating predictive models, particularly Transformers, on tasks such as per‑step transition prediction, fixed‑horizon state prediction, and sequential textual‑observation prediction. It compares byte‑level Transformer baselines with two memory‑augmented architectures, showing that a pseudo‑video spatial‑memory Transformer achieves perfect validation accuracy on selected tasks and improves sequential text‑trace prediction, while a generic auxiliary latent‑memory Transformer does not consistently help. The study highlights that structured, task‑aligned working memory can be more effective than merely increasing latent capacity, and positions the benchmark as a compact setting for testing architectures that couple textual interfaces to learned structured state.

By Alexey Potapov