Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces LOCOMO-CONV, a conversational memory benchmark that expands on the existing LoCoMo dataset with four query styles—dialog, implicit, counterfactual, and composed—designed to evaluate memory systems in realistic conversational settings. Experiments across five memory systems reveal that conversational framing uncovers significant retrieval gaps missed by traditional QA benchmarks, particularly for implicit and composed queries, and that strong retrieval does not necessarily translate into higher response quality. The study also highlights silent grounding in implicit queries, where memory enhances contextual grounding without explicitly presenting the gold fact, suggesting a need for reasoning-based memory elaboration.
arXiv:2606. 28337v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems are often evaluated using final answer accuracy, even though their failures can originate from preprocessing, retrieval, context packing, or generation.
The paper introduces a three‑layer checklist-and-judge framework to evaluate interpreter agents that mediate live conversation across languages. It assesses semantic, pragmatic, and cultural‑social dimensions—naturalness, intent, and social appropriateness—rather than just fidelity, in both single‑turn and multi‑turn settings. Extensive validation shows that conventional MT metrics miss failures in stronger interpreters, and that context, structured instructions, and cultural cues influence communicative success.
arXiv:2609.07093v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...
The paper introduces SCALE-QA, a new QA benchmark that tests conversational memory in flat, unsegmented multi‑topic threads by requiring agents to infer which earlier episode supports a later task decision. The dataset contains 3,000 audited questions across ten domains, uses deterministic four‑way multiple‑choice grading, and includes a runtime builder for reproducibility. The authors also propose Temporal‑Semantic Interleaved Memory Reconstruction (TSIM), a hierarchical memory stack that segments turns into coherent episodes and indexes them with deterministic summaries and cluster‑routing views, achieving significant accuracy gains over strong RAG baselines and long‑context LLMs.
Hy‑MultiTurn is a Chinese benchmark designed to evaluate deep multi‑turn dialogue understanding over long interactions. It introduces six controlled evaluation modes—constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution—across 209 tasks ranging from 12 to 76 turns, incorporating dialogue length, irrelevant distractions, and colloquial phrasing. Testing 22 state‑of‑the‑art models shows the benchmark is highly challenging, with even the best model meeting all criteria only 41.1% of the time and no model excelling in every mode.