Benchmarking Text Generation Inference
Related stories
Faster Text Generation with Self-Speculative Decoding
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
arXiv:2503.06868v2 Announce Type: replace-cross Abstract: Existing long-text generation methods produce lengthy outputs from short inputs, leaving long-input-to-long-output generation underexplored....
Assisted Generation: a new direction toward low-latency text generation
The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance
arXiv:2607. 24767v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relevant information to use as source material during text generation.
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
FLEET is a new method for text generation that adds a memory mechanism to large language models. It represents each generation as a sparse trajectory of high‑entropy states and uses these trajectories to compute per‑token utility scores that adjust the logits. Benchmarks show that FLEET matches the accuracy of repeated sampling while being three times faster and improving accuracy on complex coding tasks, all with minimal changes to existing pipelines.
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608. 19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.
DiffusionGemma: 4x faster text generation
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
arXiv:2601. 22546v2 Announce Type: replace-cross Abstract: The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities.
Faster Text Generation with TensorFlow and XLA
Does task decomposition improve automatic NLG evaluation?
The LLM-as-a-judge (LLMaJ) framework has emerged as a promising solution for cheap, reproducible, reference-free Natural Language Generation (NLG) evaluation. Prior work seeks to improve LLMaJ by deco...
Little Brains, Big Feats: Exploring Compact Language Models
While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.