EMA: Elastic and Performance Transparent Memory Across GPUs
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2505. 04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall.
arXiv:2609.13592v1 Announce Type: cross Abstract: GPU memory bandwidth and capacity limit throughput in large language model (LLM) inference. The GPU memory system consists of a primary tier of high-...
The paper introduces FairInference, a system that guarantees token-level latency isolation for well-behaved clients in multi-tenant LLM serving. It provides a δ-token fairness guarantee, ensuring that a token generated in isolation within time d will be produced within d + δ in a shared environment. The approach enforces per-token deadlines, bounds GPU compute sharing delays, and accounts for shared KV cache overhead, leading to reduced latency spikes and higher overall throughput compared to existing LLM serving systems.
LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. Existing solutions for performance isolatio...
arXiv:2511. 04791v2 Announce Type: replace Abstract: Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode phases.
arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.