Benchmarking Text Generation Inference
Related stories
Faster Text Generation with Self-Speculative Decoding
Assisted Generation: a new direction toward low-latency text generation
The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance
arXiv:2607. 24767v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relevant information to use as source material during text generation.
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608. 19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.
DiffusionGemma: 4x faster text generation
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
arXiv:2601. 22546v2 Announce Type: replace-cross Abstract: The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities.
Faster Text Generation with TensorFlow and XLA
Little Brains, Big Feats: Exploring Compact Language Models
While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
arXiv:2604. 13977v2 Announce Type: replace-cross Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent.
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models
arXiv:2606. 04535v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer bidirectional attention and parallel generation, enabling them to exploit global context and naturally support format-constrained tasks like parseable JSON or reasoning templates.
Little Brains, Big Feats: Exploring Compact Language Models
arXiv:2606. 30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention.