The paper introduces EAR, an Entity‑Aware Partitioning approach that improves retrieval‑augmented generation for multiple‑choice question answering by extracting normalized surface anchors from questions, answers, and the corpus. EAR retrieves local windows around matching anchors and can attach a larger parent passage via an extractive summary, reducing retrieved words by 37.5‑40.2% compared to fixed‑size chunks. Experiments on a cleaned MMLU‑style subset with Mistral, Gemma, and DeepSeek show modest accuracy changes, none statistically significant, highlighting EAR’s methodological contribution of compact, inspectable retrieval units.
By Cenab Batu Bora, Oylum Alatl{\i}, Sebnem Bora, Oguz Dikenelli
arXiv:2609.18672v1 Announce Type: new
Abstract: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual desig...
By Janghoon Lee (Redrob)
arXiv:2608. 16621v1 Announce Type: new Abstract: Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time.
By Yusuke Takahashi, Kyle Wild, Asako Uraki
Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked.
The paper argues that retrieval‑augmented question‑answering systems should perform semantic compilation at ingest time rather than re‑deriving meaning at query time. By building a maintained structure—incrementally updated embeddings and validated atomic claims—read operations become far cheaper, with experimental results showing higher accuracy and lower token usage compared to traditional chunk‑based retrieval. The authors present two proofs: cheaper incremental updates and superior performance on broadcast‑interview transcripts, suggesting a new systems agenda for compilation and read planning.
By Kyle Wild, Yusuke Takahashi, Asako Uraki
The paper presents a system for AI contact centers that answers questions only from a closed set of verified QA units, returning the unit verbatim or routing to clarification, abstention, or handoff. The index is enriched offline using staged linguistic seeding (SLS), where human-authored slot recipes are expanded by GPT‑4.1‑mini and lightly filtered by humans, enabling a single retrieval pass without query-time generation. On held‑out data from two industrial domains, SLS improves hybrid retrieval recall at rank 1 to 0.881/0.930 and outperforms doc2query by 0.20/0.32, while also reducing unsupported content from 7‑13% to near 0%.
By Hyeonseop Yoon, Jeong-Eun Park
arXiv:2608. 18752v2 Announce Type: replace-cross Abstract: Statutory retrieval is necessary for citation-grounded legal question answering, but remains underexplored for Greek.
By Ernest Beta, Odysseas S. Chlapanis, Dimitrios Galanis, Ion Androutsopoulos
Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is thereafter merely consulted -- a compiler, not an interpreter, of meaning.
arXiv:2608. 13568v1 Announce Type: cross Abstract: Coding agents spend most of their context budget on retrieval.
By Pengcheng Xu
arXiv:2608. 05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications.
By Ayoub Kirouane, Christos Petrocheilos
The paper "Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities" introduces the Waldo benchmark, a multilingual QA dataset built from Wikipedia that focuses on knowledge gaps and conflicts across languages. It evaluates eight models in five languages and finds that when a fact is missing in one language, models tend to use evidence from the other language, but when conflicting accounts exist, responses align strongly with the query language, leading to different answers for semantically identical questions. The study also explores mitigation strategies, including ablating attention heads and LoRA-based training, which can reduce the preference gap by up to 61.5%.
By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
arXiv:2605. 06647v2 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges.
By Zeyu Yang, Qi Ma, Jason Chen, Anshumali Shrivastava