arXiv Machine Learning By Xinwei Qiang, Xiang Fang, Chang Chen, Yue Guan, Yufei Ding

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Read the original on arXiv Machine Learning →

The paper introduces the concepts of an information floor and a model gap to analyze block drafting in language models. By estimating these metrics across multiple domains and models, it finds that the all-parallel floor limits per-slot acceptance to 71% on Qwen3-4B, that a single realized token can eliminate most of this floor, and that current drafters still operate far above their floors, indicating significant room for improvement. These results highlight the distinct contributions of short-range conditioning versus proposal quality in block drafting.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 27

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Block drafters generate multiple tokens in a single forward pass before earlier target tokens are produced, combining two loss components: missing within‑block path information and imperfect modeling of observable information. The study introduces an information floor—the minimum expected rejection for a given conditioning order—and defines the model gap as rejection above this floor. Across four domains and several models, the authors find that the all‑parallel floor limits per‑slot acceptance to 71% for Qwen3‑4B, that a single realized token can eliminate 86–100% of this floor, and that current drafters exhibit significant model gaps, accounting for 43–64% of DFlash rejection and 85–92% of DSpark’s oracle‑conditioned rejection.

Hugging Face Trending Papers
Jul 2

Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection.

arXiv AI
Aug 19

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.

By Javier Aguilar Mart\'in