Block drafters generate multiple tokens in a single forward pass before earlier target tokens are produced, combining two loss components: missing within‑block path information and imperfect modeling of observable information. The study introduces an information floor—the minimum expected rejection for a given conditioning order—and defines the model gap as rejection above this floor. Across four domains and several models, the authors find that the all‑parallel floor limits per‑slot acceptance to 71% for Qwen3‑4B, that a single realized token can eliminate 86–100% of this floor, and that current drafters exhibit significant model gaps, accounting for 43–64% of DFlash rejection and 85–92% of DSpark’s oracle‑conditioned rejection.
arXiv:2608.30427v1 Announce Type: cross
Abstract: Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, pres...
By Ephrem Wu
arXiv:2607. 01893v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix.
By Tianjian Yang, Meng Li
Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole block in parallel, which is fast but trained with a full-block cross-entropy that supervises every position against the gold continuation -- even though inference discards every token after the first rejection.
arXiv:2608. 10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment.
By Joyjeet Singh
The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.
By Javier Aguilar Mart\'in
LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.
By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.
By Jundong Hu, Shekar Ramachandran
arXiv:2609.23886v1 Announce Type: new
Abstract: Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears wit...
By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv:2607. 06925v1 Announce Type: new Abstract: Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block'' using a sparse set of explicit \emph{reference anchors}.
By Yufeng Wang, Lu Wei, Haibin Ling
IBBench-Light is a paired evaluation framework that tests language models on both executing procedures and reading text from the same external record. The benchmark uses twelve semantic bases to generate 144 matched pairs per model, with four instruction‑quantized models producing 1,152 greedy responses. Metrics such as Paired Exact‑Contract Accuracy (PECA) reveal that models like Qwen achieve high success on individual prompts but only 97 complete pairs, highlighting the importance of paired evaluation.
By Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao
arXiv:2607. 23191v1 Announce Type: new Abstract: Fine-tuned code LLMs can be conditioned on a lightweight design-intent header to steer parametric CAD generation, but whether the model actually reads the header's content has not been tested under a metric independent of the conditioning itself, nor with a causal control.
By Yang Xiao