arXiv AI

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

arXiv AI
Jul 31

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.

By Guanming Xiong, Penghui Zhang
arXiv AI
Sep 3

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.

By Phanindra Reddy Madduru
arXiv Machine Learning
Sep 25

RADAR: Readiness for AI Discovery and Agentic Reach

RADAR (Readiness for AI Discovery and Agentic Reach) evaluates how well AI systems serve citizens in 166 countries by testing two tasks: whether a chatbot can provide correct, officially sourced, country‑specific answers about public services (informational legibility) and whether an automated agent can actually access those services (agent operability). The study finds that AI can describe public services much better than it can reach them, with informational legibility consistently higher than agent operability across all countries and the gap remaining unchanged by national wealth. The two deficiencies have distinct causes—language representation in web corpora for legibility and national web presence for operability—and therefore require different solutions. "whyItMatters":"RADAR highlights a gap that traditional digital‑government rankings overlook, enabling governments of any income level to identify and address the specific barriers preventing AI from actually accessing public services."

By Luke Jordan, Tiago C. Peixoto, Manuel Ramos-Maqueda
arXiv AI
Sep 15

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

The paper introduces Continual Search, an iterative framework that guides large language models to persistently search for diagnostic evidence in long AI agent execution logs, addressing the limitations of one-shot judgments. Evaluated on four existing RCA benchmarks and a new large-scale dataset called MegaRCA-Mix, Continual Search consistently boosts attribution performance, achieving a 40% F1 improvement for GPT‑5.5 on MegaRCA‑Mix. The results show that effective search can outweigh raw model scale, enabling lower-tier models to outperform higher-tier ones in root‑cause attribution tasks.

By Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta
arXiv Computation and Language
Sep 11

SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs

SearchAtlas is a framework that transforms raw search trajectories of large language model (LLM) agents into structured evidential query graphs, where edges capture how evidence is propagated from queries to the final answer. The automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. Using SearchAtlas, the authors analyze five search agents on three benchmarks, uncovering systematic differences in search scale and evidence aggregation, and revealing process failures such as fragmented answer support, unmet question constraints, and unverified parametric knowledge that correlate strongly with incorrect answers.

By Jiacheng Sang, Mengyuan Li, Sanxing Chen, Yukun Huang, Yu Feng, Bhuwan Dhingra