Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference
Related stories
Assisted Generation: a new direction toward low-latency text generation
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
arXiv:2606. 14325v1 Announce Type: cross Abstract: Property Graphs are rapidly being adopted as database frameworks for representing heterogeneous data sources.
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
arXiv:2601. 22546v2 Announce Type: replace-cross Abstract: The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities.
Open-Source Text Generation & LLM Ecosystem at Hugging Face
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
arXiv:2606. 14943v1 Announce Type: cross Abstract: Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding and conditional likelihood computation.
Little Brains, Big Feats: Exploring Compact Language Models
While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608. 19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.
Faster Text Generation with TensorFlow and XLA
PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs
arXiv:2607. 20470v1 Announce Type: new Abstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets.
DiffusionGemma: 4x faster text generation
Predicting Future Behaviors in Reasoning Models Enables Better Steering
arXiv:2606. 11172v1 Announce Type: new Abstract: Deployed large reasoning models (LRMs) often behave unexpectedly.