← Back to all news
Hugging Face Blog April 16, 2025

Prefill and Decode for Concurrent Requests - Optimizing LLM Performance

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Apr 2, 2025

Efficient Request Queueing – Optimizing LLM Performance

llms
More like this →
Hugging Face Blog
Jun 12, 2025

How Long Prompts Block Other Requests - Optimizing LLM Performance

llms
More like this →
arXiv AI
Jul 3

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

arXiv:2607. 02043v1 Announce Type: cross Abstract: Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering.

By Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor R\"uhle
llmsbenchmarks
More like this →
arXiv AI
Aug 11

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

arXiv:2608. 08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control.

By Shuowei Jin, Xueshen Liu, Jiaxin Shan, Le Xu, Tieying Zhang, Liguang Xie, Z. Morley Mao
llmsefficiency
More like this →
Hugging Face Trending Papers
Jul 5

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attention prevents exact autoregressive-style KV caching, since committing one position shifts the KV activations of all others.

llmsdiffusion
More like this →
arXiv Machine Learning
Jul 7

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

arXiv:2607. 04206v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation.

By Nitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella
llmsdiffusion
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e