EnComp: Lightweight Encoder-Only Context Compression for Retrieval-Augmented Question Answering
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2505. 23277v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) often suffers from long and noisy retrieved contexts.
The paper introduces INTRA, an attention-based encoder-decoder framework that retrieves directly from its own internal representations instead of using an external retriever. By having decoder attention query pre-encoded evidence chunks, INTRA unifies retrieval and generation, eliminating the typical mismatch seen in retrieval-augmented generation pipelines. Experiments on question-answering benchmarks show that INTRA outperforms strong engineered retrieval pipelines in both evidence recall and overall answer quality.
The paper introduces Retrieval-Augmented Decoding (RAD), a decoding-time method that improves the truthfulness of large language models without retraining. RAD uses a small reference set of up to ten annotated examples to build a grounding space of context embeddings and next-token logits, which it retrieves and aggregates during inference to shape the model’s output. Experiments on four open-ended generation benchmarks and four different LLMs show that RAD consistently outperforms strong baselines and generalizes well across tasks.
arXiv:2606. 06197v1 Announce Type: cross Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs).
arXiv:2609.25537v1 Announce Type: new Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
arXiv:2606. 06906v1 Announce Type: cross Abstract: Long-context question answering (QA) remains challenging for smaller language models even when answer-bearing evidence is already present in the input.