arXiv AI

Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts

The paper introduces decision checkpoints that log observations and tool actions during inference, enabling the assignment of outcome categories to scientific-search agent responses. Using these checkpoints on 540 AutoResearchBench questions, the authors compare keyword search and read-first strategies, finding that keyword search yields 24.6% accuracy while raw search achieves 17.8%. The checkpoint protocol reveals detailed differences in target exposure, inspection attempts, and evidence-search calls that aggregate accuracy alone obscures.

arXiv AI
Aug 7

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv:2608. 05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers.

By Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
arXiv AI
Jul 8

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

arXiv:2607. 05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect.

By Yufeng Wang
arXiv AI
Aug 24

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.

By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
arXiv Computation and Language
Sep 11

SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs

SearchAtlas is a framework that transforms raw search trajectories of large language model (LLM) agents into structured evidential query graphs, where edges capture how evidence is propagated from queries to the final answer. The automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. Using SearchAtlas, the authors analyze five search agents on three benchmarks, uncovering systematic differences in search scale and evidence aggregation, and revealing process failures such as fragmented answer support, unmet question constraints, and unverified parametric knowledge that correlate strongly with incorrect answers.

By Jiacheng Sang, Mengyuan Li, Sanxing Chen, Yukun Huang, Yu Feng, Bhuwan Dhingra
arXiv AI
Sep 15

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

The paper introduces Continual Search, an iterative framework that guides large language models to persistently search for diagnostic evidence in long AI agent execution logs, addressing the limitations of one-shot judgments. Evaluated on four existing RCA benchmarks and a new large-scale dataset called MegaRCA-Mix, Continual Search consistently boosts attribution performance, achieving a 40% F1 improvement for GPT‑5.5 on MegaRCA‑Mix. The results show that effective search can outweigh raw model scale, enabling lower-tier models to outperform higher-tier ones in root‑cause attribution tasks.

By Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta