arXiv AI

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

arXiv:2607. 20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs.

arXiv AI
Sep 24

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

FLEET is a new method for text generation that adds a memory mechanism to large language models. It represents each generation as a sparse trajectory of high‑entropy states and uses these trajectories to compute per‑token utility scores that adjust the logits. Benchmarks show that FLEET matches the accuracy of repeated sampling while being three times faster and improving accuracy on complex coding tasks, all with minimal changes to existing pipelines.

By Oleksii Streltsov, Oleksandra Vitko
arXiv Computation and Language
Sep 4

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.

By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
Hugging Face Trending Papers
6d ago

Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs

Masked diffusion language models can generate outputs in any order, but recent findings suggest this flexibility may reduce diversity by delaying uncertain tokens. The study identifies low‑confidence remasking (LCR) as the main culprit, showing that its filtering of lower‑probability tokens suppresses diversity exponentially. Replacing LCR with top‑probability position selection (TPP) restores diversity, and adding Entropy‑Guided Initialization (EGI) further enhances rollout diversity and solution coverage, demonstrating the benefits of arbitrary‑order generation for diverse outputs.

arXiv AI
Aug 26

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec is a framework that improves speculative decoding for large language models by reshaping the residual distribution during verification. It addresses the problem of residual drift, where rejected candidates cause the target distribution to diverge from the draft model’s predictions, rendering later candidates ineffective. By aligning the verification process with the draft model’s high‑confidence regions, ResiSpec prevents candidate obsolescence and achieves up to 1.92× speedup over existing multi‑candidate methods.

By Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
arXiv AI
Aug 28

Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models

The paper introduces Dual-Seed Comparison (DSC), a protocol that uses two independent LLM-generated seeds to reduce systematic bias in probabilistic sampling. DSC constructs a bit sequence from the character-level ordinal values of the seeds, normalizes it into a pseudo-uniform variate, and maps it to the target distribution via the inverse cumulative distribution function. Empirical results show DSC outperforms existing methods in 96% of evaluated settings and enhances distributional control in tasks like MCQ generation and attribute-constrained text-to-image prompting.

By Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui