SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification.
arXiv:2606. 12243v1 Announce Type: cross Abstract: Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel.
arXiv:2606. 01019v1 Announce Type: cross Abstract: Large Language Model (LLM) generation remains expensive because autoregressive decoding calls the model once for each new token.
arXiv:2603. 18016v2 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model.
The paper introduces DLoop, a looped speculative decoding technique that adaptively performs multiple drafting stages before verification, allowing a draft model to continue generating tokens while confident. By verifying all accumulated draft tokens together and training the draft model to handle its own hidden states for unverified tokens, DLoop reduces the number of target‑model forward passes needed. Experiments across several speculative decoding methods show wall‑clock speedups of 5–41 % without sacrificing lossless decoding.
arXiv:2607. 03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns.