arXiv:2607. 06523v1 Announce Type: new Abstract: Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation.
By Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera
arXiv:2608.00311v2 Announce Type: replace
Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and th...
By Maryam Haghifam, Jason Cong, Yizhou Sun
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed.
The paper introduces Ladder Side Tuning (LST), a parameter‑efficient fine‑tuning method that adds a lightweight side network to large language models. LST matches QLoRA’s compute scaling while halving peak memory usage, enabling 7B‑parameter models to be fine‑tuned on a single 12 GB GPU with 2k‑token contexts without gradient checkpointing. The authors also present xLadder, a depth‑extended variant that increases effective depth through cross‑connections, allowing deeper reasoning without extra memory overhead.
By Estelle Zheng, Nathan Cerisara, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara
arXiv:2607. 07388v1 Announce Type: cross Abstract: Large Language Models (LLMs) store factual knowledge and domain-specific patterns implicitly in dense Transformer parameters, making knowledge expansion costly through pretraining, fine-tuning, retrieval augmentation, or longer contexts.
By Yutang Ma, Kecheng Huang, Xikun Jiang, Zili Shao
Lngram v2 introduces a latent N‑gram memory system that decouples memory routes, memory dimension, and backbone width, enabling scalable memory capacity for transformers. It employs context‑aware grouped‑query attention, a zero‑value sink, and counterfactual surrogate gradients to improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments on vision‑language models up to 30B parameters show consistent performance gains, reduced memory parameters, and stable semantic structure in the discrete IDs.
By Yunao Zheng, Bin Wen, Xiaojie Wang
The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.
By Eduardo Novaes Hering
arXiv:2609.25537v1 Announce Type: new
Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
By Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu
arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.
By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
arXiv:2606. 29563v1 Announce Type: cross Abstract: Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs.
By Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi, Golnoosh Samei
arXiv:2609.22100v1 Announce Type: cross
Abstract: Retrieval-augmented generation (RAG) improves language models with retrieved evidence, but processing many long passages is costly and can introduce...
By Artem Sakhno, Grigorii Davydenko, Omar Zoloev, Julia Belikova, Andrey Savchenko, Maksim Makarenko
The paper investigates how to effectively pre‑train language models when the data budget is limited but compute is plentiful. It shows that increasing model size only improves performance up to an optimal point, after which overfitting degrades generalization, and that this optimal size varies with both the data budget and downstream tasks. To overcome the inefficiencies of standard Transformers in this regime, the authors propose recursive Transformers that reuse a shared block across depth and employ factorized embeddings, achieving better results than standard models on 10M–100M word pre‑training budgets and competitive performance with BabyLM Challenge 2025 winners.
By Serdar G\"ulbahar, Lukas Edman, Alexander Fraser