arXiv:2602. 22221v2 Announce Type: replace-cross Abstract: Search engines and AI-powered systems increasingly mediate access to factual information, yet their reliability remains difficult to evaluate in realistic information-seeking settings.
By Geng Liu, Li Feng, Mengxiao Zhu, Francesco Pierri
arXiv:2606. 02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis.
By Jiaming Wang, Ziteng Feng, Jiangtao Wu, Ruihao Li, Qianqian Xie, Yuxiang Ren, He Zhu, Xueming Han, Fanyu Meng, Junlan Feng, Jiaheng Liu
arXiv:2605.29224v2 Announce Type: replace-cross
Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incor...
By Aditya Nawal, Manit Baser, Mohan Gurusamy
arXiv:2606. 05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference.
By Yongjie Wang, Xinyue Zhang, Kunhong Yao, Zhiwei Zeng, Kaisong Song, Jun Lin, Zhiqi Shen
arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.
By Guanming Xiong, Penghui Zhang
The paper reports a case study of a large language model (LLM) coding agent tasked with building a multi‑component data system from a detailed specification. During a single session the agent introduced five defects, which were categorized by violated constraints and detection methods. The study also evaluates the agent’s retrieval‑filtering strategy on the HotpotQA benchmark, showing that filtering to a graph‑identified entity set yields higher recall than unfiltered search, with a statistically significant gap across all tested budgets.
By Phanindra Reddy Madduru
RADAR (Readiness for AI Discovery and Agentic Reach) evaluates how well AI systems serve citizens in 166 countries by testing two tasks: whether a chatbot can provide correct, officially sourced, country‑specific answers about public services (informational legibility) and whether an automated agent can actually access those services (agent operability). The study finds that AI can describe public services much better than it can reach them, with informational legibility consistently higher than agent operability across all countries and the gap remaining unchanged by national wealth. The two deficiencies have distinct causes—language representation in web corpora for legibility and national web presence for operability—and therefore require different solutions.
"whyItMatters":"RADAR highlights a gap that traditional digital‑government rankings overlook, enabling governments of any income level to identify and address the specific barriers preventing AI from actually accessing public services."
By Luke Jordan, Tiago C. Peixoto, Manuel Ramos-Maqueda
The paper introduces Continual Search, an iterative framework that guides large language models to persistently search for diagnostic evidence in long AI agent execution logs, addressing the limitations of one-shot judgments. Evaluated on four existing RCA benchmarks and a new large-scale dataset called MegaRCA-Mix, Continual Search consistently boosts attribution performance, achieving a 40% F1 improvement for GPT‑5.5 on MegaRCA‑Mix. The results show that effective search can outweigh raw model scale, enabling lower-tier models to outperform higher-tier ones in root‑cause attribution tasks.
By Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta
arXiv:2509. 00761v4 Announce Type: replace Abstract: Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy.
By Boqin Yuan, Ziqi Wang
arXiv:2603. 00801v2 Announce Type: replace Abstract: Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources.
By Shrey Shah, Levent Ozgur
arXiv:2605. 06647v2 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges.
By Zeyu Yang, Qi Ma, Jason Chen, Anshumali Shrivastava
SearchAtlas is a framework that transforms raw search trajectories of large language model (LLM) agents into structured evidential query graphs, where edges capture how evidence is propagated from queries to the final answer. The automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. Using SearchAtlas, the authors analyze five search agents on three benchmarks, uncovering systematic differences in search scale and evidence aggregation, and revealing process failures such as fragmented answer support, unmet question constraints, and unverified parametric knowledge that correlate strongly with incorrect answers.
By Jiacheng Sang, Mengyuan Li, Sanxing Chen, Yukun Huang, Yu Feng, Bhuwan Dhingra