Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
arXiv:2602. 22647v2 Announce Type: replace-cross Abstract: Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation.
arXiv:2607. 10044v1 Announce Type: new Abstract: Constrained decoding is essential in generative retrieval, where document identifiers generated directly from a query must exactly match a predefined library of valid IDs.
arXiv:2602. 22647v2 Announce Type: replace-cross Abstract: Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation.
arXiv:2601. 07048v5 Announce Type: replace-cross Abstract: Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications.
arXiv:2609.05760v1 Announce Type: cross Abstract: We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU enviro...
Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.
arXiv:2609.21281v1 Announce Type: cross Abstract: Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, exp...
The paper introduces the Tri‑Metric Router, a deterministic, training‑free policy that chooses among Raw, Neural, and Lexical pipelines for retrieval‑augmented generation on commodity GPUs. It uses three CPU‑side signals—spatial complexity, syntactic density, and type‑token ratio—to balance VRAM headroom and latency, calibrated on LongBench qasper. The method eliminates out‑of‑memory failures and improves alignment and F1 scores compared to always‑on lexical compression without extra VRAM or training costs.
arXiv:2607. 00760v1 Announce Type: new Abstract: Long-context LLM services now sustain prompts with hundreds of thousands to millions of tokens, making the key-value (KV) cache a first-order serving cost.
arXiv:2609.38090v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-...
The paper introduces Faster Flash Decoding (FFD), a hardware‑algorithm co‑design that fuses the selector and computation into a single kernel to eliminate memory‑bandwidth bottlenecks in long‑context decoding. By replacing metadata indices with low‑bit quantized, content‑aware scanning and employing a top‑delta strategy for dynamic block filtering, FFD achieves up to 11.6× kernel‑level speedup and scales to 256K context length while preserving model accuracy. The approach is training‑free, plug‑and‑play, and demonstrates significant throughput gains on benchmarks such as RULER and LongBench.
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.
FlashAttention-V is a blocked FlashAttention implementation optimized for scalable vector architectures, designed to reduce the memory bandwidth bottleneck of transformer attention on CPUs. By fusing operations, exploiting parallelism across attention heads, and inter‑head packing, it improves vector register utilization and memory locality, enabling efficient scaling from short to very long vectors. Benchmarks on TinyLlama, Llama 3.2, Qwen2.5, and Pythia‑410M show 22×–42× speedups over scalar FlashAttention in prefill and 8×–11× in decode on a Banana Pi BPI‑F3, while also revealing quantization‑related bottlenecks that limit long‑vector scalability.
arXiv:2606. 09079v1 Announce Type: cross Abstract: Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving.