FLEET is a new method for text generation that adds a memory mechanism to large language models. It represents each generation as a sparse trajectory of high‑entropy states and uses these trajectories to compute per‑token utility scores that adjust the logits. Benchmarks show that FLEET matches the accuracy of repeated sampling while being three times faster and improving accuracy on complex coding tasks, all with minimal changes to existing pipelines.
By Oleksii Streltsov, Oleksandra Vitko
The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.
By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
arXiv:2609.24150v1 Announce Type: new
Abstract: Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are...
By Tianhua Xia, Mugilan Ganesan, Yifei Feng, Haiyu Wang, Maximilian Egger, Sai Qian Zhang
arXiv:2609.16450v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative...
By Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, Haoran You
arXiv:2610.01687v1 Announce Type: cross
Abstract: Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates ever...
By Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza
arXiv:2609.39560v1 Announce Type: new
Abstract: Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional...
By Michael Helcig, Martin Jaggi
arXiv:2608. 07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated.
By Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou
arXiv:2606. 08048v1 Announce Type: cross Abstract: Diffusion language models (DLMs) offer substantial speed advantages through parallel decoding, but the lack of token dependencies limits generation quality compared to autoregressive (AR) models.
By Juntong Shi, Brian L. Trippe, Jure Leskovec, Stefano Ermon, Minkai Xu
arXiv:2512.23765v2 Announce Type: replace-cross
Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens and a stronge...
By Tiancheng Su, Meicong Zhang, Guoxiu He
Masked diffusion language models can generate outputs in any order, but recent findings suggest this flexibility may reduce diversity by delaying uncertain tokens. The study identifies low‑confidence remasking (LCR) as the main culprit, showing that its filtering of lower‑probability tokens suppresses diversity exponentially. Replacing LCR with top‑probability position selection (TPP) restores diversity, and adding Entropy‑Guided Initialization (EGI) further enhances rollout diversity and solution coverage, demonstrating the benefits of arbitrary‑order generation for diverse outputs.
ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.
By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
The paper introduces Dual-Seed Comparison (DSC), a protocol that uses two independent LLM-generated seeds to reduce systematic bias in probabilistic sampling. DSC constructs a bit sequence from the character-level ordinal values of the seeds, normalizes it into a pseudo-uniform variate, and maps it to the target distribution via the inverse cumulative distribution function. Empirical results show DSC outperforms existing methods in 96% of evaluated settings and enhances distributional control in tasks like MCQ generation and attribute-constrained text-to-image prompting.
By Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui