arXiv AI By Xuezhen Xie, Zhiqiang Zhou

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference

Read the original on arXiv AI →

arXiv:2606. 10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 3

Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits

arXiv:2511. 11346v2 Announce Type: replace Abstract: Multi-token prediction (MTP) is a prominent strategy to significantly speed up generation in large language models (LLMs), especially in byte-level LLMs, which are tokeniser-free but prohibitively slow.

By Andreas Grivas, Lorenzo Loconte, Emile van Krieken, Piotr Nawrot, Yu Zhao, Euan Wielewski, Pasquale Minervini, Edoardo Ponti, Antonio Vergari