arXiv Machine Learning By Xin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo, Tianyi Zhou, Shenji Wan

InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context

Read the original on arXiv Machine Learning →

arXiv:2603. 05353v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) for long-context question answering is bottlenecked by inference-time prefilling over large retrieved contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
4d ago

Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

The paper introduces SCSP, a training‑free framework that improves long‑context embeddings by selectively pooling informative tokens. SCSP partitions documents into sentence‑aware chunks, adds a semantic compression prompt to each chunk, and uses prompt‑isolated attention masks to estimate token importance. The selected tokens’ intermediate‑layer representations are aggregated to form the final embedding, yielding consistent performance gains across zero‑shot and fine‑tuned models on long‑context benchmarks.

By Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu
arXiv Machine Learning
1d ago

Retrieval from Within: An Intrinsic Capability of Attention-Based Models

The paper introduces INTRA, an attention-based encoder-decoder framework that retrieves directly from its own internal representations instead of using an external retriever. By having decoder attention query pre-encoded evidence chunks, INTRA unifies retrieval and generation, eliminating the typical mismatch seen in retrieval-augmented generation pipelines. Experiments on question-answering benchmarks show that INTRA outperforms strong engineered retrieval pipelines in both evidence recall and overall answer quality.

By Elad Hoffer, Yochai Blau, Edan Kinderman, Ron Banner, Daniel Soudry, Boris Ginsburg
Hugging Face Trending Papers
Jun 4

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the prefill stage a dominant serving cost. RAG cache fusion reduces this cost by reusing precomputed key-value (KV) caches for retrieved chunks and selectively recomputing tokens under the current prompt.