arXiv AI By Xuezhen Xie, Zhiqiang Zhou

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference

Read the original on arXiv AI →

arXiv:2606. 10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 3

Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits

arXiv:2511. 11346v2 Announce Type: replace Abstract: Multi-token prediction (MTP) is a prominent strategy to significantly speed up generation in large language models (LLMs), especially in byte-level LLMs, which are tokeniser-free but prohibitively slow.

By Andreas Grivas, Lorenzo Loconte, Emile van Krieken, Piotr Nawrot, Yu Zhao, Euan Wielewski, Pasquale Minervini, Edoardo Ponti, Antonio Vergari
arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo
arXiv AI
2d ago

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

arXiv:2610.00888v1 Announce Type: cross Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per fo...

By Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song