arXiv AI

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

The study compares two methods for grounding assistants in a small Greek–English agricultural knowledge base: tool‑calling retrieval via a live data interface and vector retrieval‑augmented generation (RAG). Using the KyGround benchmark of 198 questions, vector RAG achieved 95.3% accuracy on canonical Greek questions, outperforming the tool agent’s 71.6% and revealing that the tool agent’s failures stem from literal searches that miss non‑verbatim matches. The results show that search tolerance to user typing variations—such as accents, capitalization, and Greeklish transliterations—is crucial for reliable community knowledge interfaces.

arXiv Computation and Language
Sep 14

EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development

The paper introduces EAR, an Entity‑Aware Partitioning approach that improves retrieval‑augmented generation for multiple‑choice question answering by extracting normalized surface anchors from questions, answers, and the corpus. EAR retrieves local windows around matching anchors and can attach a larger parent passage via an extractive summary, reducing retrieved words by 37.5‑40.2% compared to fixed‑size chunks. Experiments on a cleaned MMLU‑style subset with Mistral, Gemma, and DeepSeek show modest accuracy changes, none statistically significant, highlighting EAR’s methodological contribution of compact, inspectable retrieval units.

By Cenab Batu Bora, Oylum Alatl{\i}, Sebnem Bora, Oguz Dikenelli
Hugging Face Trending Papers
Aug 10

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked.

arXiv AI
Aug 24

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

The paper argues that retrieval‑augmented question‑answering systems should perform semantic compilation at ingest time rather than re‑deriving meaning at query time. By building a maintained structure—incrementally updated embeddings and validated atomic claims—read operations become far cheaper, with experimental results showing higher accuracy and lower token usage compared to traditional chunk‑based retrieval. The authors present two proofs: cheaper incremental updates and superior performance on broadcast‑interview transcripts, suggesting a new systems agenda for compilation and read planning.

By Kyle Wild, Yusuke Takahashi, Asako Uraki
arXiv Computation and Language
Sep 2

Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers

The paper presents a system for AI contact centers that answers questions only from a closed set of verified QA units, returning the unit verbatim or routing to clarification, abstention, or handoff. The index is enriched offline using staged linguistic seeding (SLS), where human-authored slot recipes are expanded by GPT‑4.1‑mini and lightly filtered by humans, enabling a single retrieval pass without query-time generation. On held‑out data from two industrial domains, SLS improves hybrid retrieval recall at rank 1 to 0.881/0.930 and outperforms doc2query by 0.20/0.32, while also reducing unsupported content from 7‑13% to near 0%.

By Hyeonseop Yoon, Jeong-Eun Park
arXiv AI
Aug 6

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

arXiv:2608. 05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications.

By Ayoub Kirouane, Christos Petrocheilos
arXiv AI
6d ago

Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

The paper "Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities" introduces the Waldo benchmark, a multilingual QA dataset built from Wikipedia that focuses on knowledge gaps and conflicts across languages. It evaluates eight models in five languages and finds that when a fact is missing in one language, models tend to use evidence from the other language, but when conflicting accounts exist, responses align strongly with the query language, leading to different answers for semantically identical questions. The study also explores mitigation strategies, including ablating attention heads and LoRA-based training, which can reduce the preference gap by up to 61.5%.

By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao