← Back to all news
arXiv Machine Learning September 1, 2026 By Daeha Lee, Do-Hyung Kim, Jae-Hong Kim

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 29

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

arXiv:2605. 06675v2 Announce Type: replace Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottleneck for serving.

By Fei Zuo, Zikang Zhou, Hao Cong, Xiaoyan Xi, Ho Fai Leung
llmsefficiency
More like this →
arXiv Machine Learning
Jul 3

Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference

arXiv:2607. 01831v1 Announce Type: cross Abstract: Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems.

By Wenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker
llmsragagentsefficiencybenchmarks
More like this →
arXiv Machine Learning
Jul 2

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache

arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.

By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung
llmsefficiencysafety
More like this →
arXiv AI
Aug 25

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

arXiv:2608.21362v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing...

By Srihari Unnikrishnan
llmsefficiency
More like this →
arXiv Machine Learning
Jul 20

VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs

arXiv:2607. 15498v1 Announce Type: cross Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference.

By Shahrzad Esmat, Dhawal Shah, Ali Jannesari
llmsefficiencybenchmarks
More like this →
arXiv AI
Jun 6

Channel-Wise Mixed-Precision Quantization for Large Language Models

arXiv:2410. 13056v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remains challenging due to the substantial memory requirements imposed by their large parameter sizes.

By Zihan Chen, Bike Xie, Jundong Li, Cong Shen
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea