Hugging Face Blog

Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference

arXiv AI
Aug 24

SCOPE: A Generative Approach for LLM Prompt Compression

SCOPE is a training‑free generative prompt‑compression framework that reduces LLM input length by chunking a prompt into semantically coherent segments, rewriting each chunk to be more concise, and then reconstructing a coherent prompt. Unlike token‑removal methods, SCOPE’s chunk‑level rewriting preserves critical information and text coherence, and includes optimization techniques for finer‑grained control of compression ratios. Extensive evaluations on question‑answering and summarization tasks show that SCOPE consistently outperforms selective compression baselines, especially at high compression ratios.

By Tinghui Zhang, Yifan Wang, Daisy Zhe Wang
arXiv AI
Sep 24

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

FLEET is a new method for text generation that adds a memory mechanism to large language models. It represents each generation as a sparse trajectory of high‑entropy states and uses these trajectories to compute per‑token utility scores that adjust the logits. Benchmarks show that FLEET matches the accuracy of repeated sampling while being three times faster and improving accuracy on complex coding tasks, all with minimal changes to existing pipelines.

By Oleksii Streltsov, Oleksandra Vitko
Hugging Face Trending Papers
Jun 29

Little Brains, Big Feats: Exploring Compact Language Models

While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains; yet, they receive far less attention. In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.

arXiv Machine Learning
Sep 23

The Probabilistic Structure of Large Language Models

The paper offers a unified probabilistic framework for large language models, describing them as probability measures over token sequences defined by autoregressive conditional distributions. Training is cast as maximum‑likelihood estimation solved via stochastic gradient methods, while generation is treated as sequential simulation of the resulting stochastic process. It also explores how the asymmetry of the Kullback–Leibler divergence relates to hallucination and the distinction between plausibility and truth, and extends the perspective to diffusion models that generate data by simulating a reverse‑time stochastic process.

By Adnan Aboulala\^a