The paper introduces the concepts of an information floor and a model gap to analyze block drafting in language models. By estimating these metrics across multiple domains and models, it finds that the all-parallel floor limits per-slot acceptance to 71% on Qwen3-4B, that a single realized token can eliminate most of this floor, and that current drafters still operate far above their floors, indicating significant room for improvement. These results highlight the distinct contributions of short-range conditioning versus proposal quality in block drafting.
By Xinwei Qiang, Xiang Fang, Chang Chen, Yue Guan, Yufei Ding
arXiv:2608.30427v1 Announce Type: cross
Abstract: Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, pres...
By Ephrem Wu
arXiv:2607. 01893v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix.
By Tianjian Yang, Meng Li
Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection.
LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.
By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
The study evaluates whether prompt‑token counts can reliably identify the lineage of large language models served via APIs. Using a frozen‑threshold approach on 24 labeled endpoint pairs, the authors find that token‑count consistency perfectly separates development pairs but only half of the holdout pairs meet the strict repeatability criteria, yielding moderate accuracy and perfect specificity. The results confirm token‑count consistency as a fingerprint of shared tokenization stacks but reject it as a standalone test for model‑family attribution.
By Bo Chen