arXiv Machine Learning

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

The paper introduces the concepts of an information floor and a model gap to analyze block drafting in language models. By estimating these metrics across multiple domains and models, it finds that the all-parallel floor limits per-slot acceptance to 71% on Qwen3-4B, that a single realized token can eliminate most of this floor, and that current drafters still operate far above their floors, indicating significant room for improvement. These results highlight the distinct contributions of short-range conditioning versus proposal quality in block drafting.

Hugging Face Trending Papers
Aug 27

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Block drafters generate multiple tokens in a single forward pass before earlier target tokens are produced, combining two loss components: missing within‑block path information and imperfect modeling of observable information. The study introduces an information floor—the minimum expected rejection for a given conditioning order—and defines the model gap as rejection above this floor. Across four domains and several models, the authors find that the all‑parallel floor limits per‑slot acceptance to 71% for Qwen3‑4B, that a single realized token can eliminate 86–100% of this floor, and that current drafters exhibit significant model gaps, accounting for 43–64% of DFlash rejection and 85–92% of DSpark’s oracle‑conditioned rejection.

Hugging Face Trending Papers
Jul 2

Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection.

arXiv AI
Aug 19

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.

By Javier Aguilar Mart\'in
arXiv AI
6d ago

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.

By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
arXiv AI
2d ago

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.

By Jundong Hu, Shekar Ramachandran
arXiv AI
Sep 15

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

IBBench-Light is a paired evaluation framework that tests language models on both executing procedures and reading text from the same external record. The benchmark uses twelve semantic bases to generate 144 matched pairs per model, with four instruction‑quantized models producing 1,152 greedy responses. Metrics such as Paired Exact‑Contract Accuracy (PECA) reveal that models like Qwen achieve high success on individual prompts but only 97 complete pairs, highlighting the importance of paired evaluation.

By Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao
arXiv Machine Learning
Jul 28

Wrong Design Intent Is Worse Than None: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion

arXiv:2607. 23191v1 Announce Type: new Abstract: Fine-tuned code LLMs can be conditioned on a lightweight design-intent header to steer parametric CAD generation, but whether the model actually reads the header's content has not been tested under a metric independent of the conditioning itself, nor with a causal control.

By Yang Xiao