FLEET is a new method for text generation that adds a memory mechanism to large language models. It represents each generation as a sparse trajectory of high‑entropy states and uses these trajectories to compute per‑token utility scores that adjust the logits. Benchmarks show that FLEET matches the accuracy of repeated sampling while being three times faster and improving accuracy on complex coding tasks, all with minimal changes to existing pipelines.
By Oleksii Streltsov, Oleksandra Vitko
The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.
By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
arXiv:2609.24150v1 Announce Type: new
Abstract: Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are...
By Tianhua Xia, Mugilan Ganesan, Yifei Feng, Haiyu Wang, Maximilian Egger, Sai Qian Zhang
arXiv:2609.16450v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative...
By Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, Haoran You
arXiv:2610.01687v1 Announce Type: cross
Abstract: Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates ever...
By Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza
arXiv:2609.39560v1 Announce Type: new
Abstract: Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional...
By Michael Helcig, Martin Jaggi