arXiv Machine Learning

TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories

arXiv AI
Jul 3

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning

arXiv:2607. 01241v1 Announce Type: cross Abstract: Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is often spread across multiple locations and connected through both local syntactic dependencies and global semantic relations.

By Yaxin Gao, Yao Lu, Jinhong Deng, Jiaqi Nie, Zhe Tang, Jian Zhang, Zhaowei Zhu, Shanqing Yu, Qi Xuan, Joey Tianyi Zhou
arXiv Machine Learning
1d ago

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

CommunityKV is a new framework that treats sparse attention as a community detection problem, building a token graph from $QK^T$ scores and partitioning it into semantically coherent communities. It updates token communities in constant time during streaming decoding, avoiding costly global re‑partitioning. Experiments on Qwen3 and Llama‑3.1 show that CommunityKV can increase end‑to‑end generation throughput by up to 1.25×, and with query‑group graph aggregation up to 1.71×, while maintaining comparable accuracy.

By Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu
arXiv Computation and Language
4d ago

Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

The paper introduces SCSP, a training‑free framework that improves long‑context embeddings by selectively pooling informative tokens. SCSP partitions documents into sentence‑aware chunks, adds a semantic compression prompt to each chunk, and uses prompt‑isolated attention masks to estimate token importance. The selected tokens’ intermediate‑layer representations are aggregated to form the final embedding, yielding consistent performance gains across zero‑shot and fine‑tuned models on long‑context benchmarks.

By Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu
arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
arXiv Computation and Language
Sep 7

CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation

CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation introduces a reranking framework that evaluates and enhances the coherence of retrieved passages across four dimensions—Intra-Domain Relevance, Noise Resistance, Informational Bonding, and Factual Consistency. The method transforms passages into directed heterogeneous entity graphs, reweights factual anchors, encodes structural patterns with a Relational Graph Convolutional Network, and fuses inter-chunk coherence with query relevance to produce a final ranking. Evaluations on four multi‑hop benchmarks show that CAGE matches or surpasses strong baselines, improving Recall@5 on bridge‑dominated datasets and consistently boosting downstream Exact Match scores, indicating that structurally coherent context leads to more precise answers even when retrieval recall is similar or lower.

By Tong Qi, Jingyu Wu, Youbing Yin, Spencer Hong, Daben Liu, Erin Babinsky
Hugging Face Trending Papers
Jul 9

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness.