CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 22221v2 Announce Type: replace-cross Abstract: Search engines and AI-powered systems increasingly mediate access to factual information, yet their reliability remains difficult to evaluate in realistic information-seeking settings.
arXiv:2606. 02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis.
arXiv:2605.29224v2 Announce Type: replace-cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incor...
arXiv:2606. 05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference.
arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.
The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.