arXiv AI

Length-Adaptive Decoding for Masked Diffusion Machine Translation

The paper introduces Entropy-Valley (EV), a training‑free method for selecting target length in masked diffusion machine translation. EV evaluates candidate canvases by mean predictive entropy from all‑mask forward passes, choosing the length the model is best prepared to fill. Compared to a baseline that uses training‑corpus length statistics, EV recovers a substantial portion of the COMET‑22 gain across En→Zh, Zh→En, and En→De, and expert evaluation confirms adequacy improvements, especially for Zh→En.

arXiv Machine Learning
Jul 7

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.

By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
arXiv AI
Sep 3

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

The paper introduces PILL, a new infilling technique for diffusion language models that eliminates the need for a preset initial length and reduces inference overhead. PILL uses probing-based length-free decoding, cutting down on extra forward passes and speeding up generation. Experiments across five diffusion models and eight benchmarks show PILL outperforms the strongest baseline with higher pass rates and BLEU-2 scores while running 1.82× faster.

By Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong
arXiv AI
Jun 10

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv:2606. 10829v1 Announce Type: cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled.

By Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
arXiv Computation and Language
Aug 28

Survival-Guided Length Control for Efficient Diffusion Language Models

The paper introduces a survival-guided length control method for diffusion language models (DLMs), framing length selection as a discrete-time survival problem over the end-of-sequence token. This training‑free, plug‑in length predictor can be added to any existing DLM and reduces unnecessary denoising steps. Experiments on reasoning and code‑generation benchmarks show up to a seven‑fold speedup in inference while maintaining task accuracy, and reveal that predicted lengths vary significantly even within the same dataset, affecting model performance.

By Ivan Kobyzev, Abbas Ghaddar, Yufei Cui
Hugging Face Trending Papers
Jul 6

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains.