arXiv AI

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents

arXiv:2608. 02751v2 Announce Type: replace-cross Abstract: Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata.

arXiv Computation and Language
Aug 31

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents

The paper introduces Sieve, a search‑inspect‑fetch framework that leverages a Boolean Query Language (BQL) to target specific webpage fields, rank candidates, present structure‑rich result cards, and fetch only selected sections. Compared to traditional Search‑Visit agents, Sieve achieves higher accuracy across three QA collections while reducing token usage by 20.7–50.6%. Boolean filtering consistently improves performance for all tested rankers and remains effective across different retrievers and agent backbones.

By Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
Hugging Face Trending Papers
Sep 8

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.

arXiv AI
Jul 31

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.

By Guanming Xiong, Penghui Zhang
arXiv Computation and Language
Sep 24

Improving LLM-based Autonomous Web Agents with Filtering

The paper investigates how to improve large language model (LLM) based autonomous web agents by filtering irrelevant webpage content. The authors reproduce baseline models on the WebArena benchmark and identify failure modes caused by raw HTML input. They propose DeBERTa‑ and T5‑based retrieval models that rank HTML elements by relevance, fine‑tuned on Mind2Web data, and demonstrate that the DeBERTa model raises the LLaMA‑2‑70B agent’s success rate from 1.97% to 2.96%. Additionally, a zero‑shot ColBERT retriever achieves recall of 0.52 on Mind2Web and 0.47 on WebArena.

By Zhitong Guo, Jing Yu Koh, Ruiyu Li
arXiv AI
Aug 24

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.

By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
arXiv AI
3d ago

TRACE: Trajectory Selection for Parallel Scaling of Search Agents

TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.

By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan