Hugging Face Trending Papers

The Optimization Landscape of Learning Compacted Context Models

Read the original on Hugging Face Trending Papers →

The paper investigates the challenges of optimizing compacted context models, particularly the KV cache, in continual learning scenarios. It identifies the optimization landscape as brittle and flat, and proposes a simplified Perceiver-based architecture that matches or surpasses full Perceiver transformers in continuous context compaction. Experiments on MCQ tasks in Finance, Legal, Gutenberg, and Code demonstrate the effectiveness of this approach.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
Jun 8

End-to-End Context Compression at Scale

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.

arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov