arXiv Computation and Language By Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa, Yoshimasa Tsuruoka

FlexComp: One Model for Every Ratio in Context Compression

Read the original on arXiv Computation and Language →

FlexComp is a framework that allows a single model to perform context compression at any desired ratio, unlike existing methods that require separate models for each fixed ratio. It achieves this by sampling a memory budget during training and selecting the appropriate budget at inference time using either confidence-based cascade routing or a lightweight learned predictor. Experiments on ICAE, 500xCompressor, and SAC show that FlexComp matches the performance of specialized fixed-ratio models while enabling high compression rates and improving decoding throughput.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
arXiv Machine Learning
Sep 17

Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs

The paper introduces the Tri‑Metric Router, a deterministic, training‑free policy that chooses among Raw, Neural, and Lexical pipelines for retrieval‑augmented generation on commodity GPUs. It uses three CPU‑side signals—spatial complexity, syntactic density, and type‑token ratio—to balance VRAM headroom and latency, calibrated on LongBench qasper. The method eliminates out‑of‑memory failures and improves alignment and F1 scores compared to always‑on lexical compression without extra VRAM or training costs.

By Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra, Shashank Kapadia, Haoan Wang, Anupriya Sharma
Hugging Face Trending Papers
Jun 8

End-to-End Context Compression at Scale

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.