WebArxiv is a reproducible benchmark designed to evaluate multimodal web agents on arXiv-related tasks. It consists of 510 static, time‑invariant tasks that require multi‑constraint paper retrieval, fine‑grained content extraction, and cross‑paper comparison, each with a deterministic ground truth. The benchmark highlights challenges for foundation‑model agents, such as over‑reliance on fixed interaction histories, and introduces a lightweight dynamic‑memory mechanism to improve adaptive retrieval and reasoning.
By Zihao Sun, Zijing Shi, Ling Chen
arXiv:2508.13186v2 Announce Type: replace-cross
Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However...
By Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Hui Huang, Donghao Zhou, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Zhaoxiang Zhang, Qiangpeng Yang, Shilei Wen
arXiv:2606.13120v2 Announce Type: replace
Abstract: Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing be...
By Yunhan Wang, Jiaan Wang, Lianzhe Huang, Xianfeng Zeng, Fandong Meng
We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own search-and-verification process rather than scaling data, tools, or auxiliary model components.
arXiv:2609.14412v1 Announce Type: new
Abstract: Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Recent work on reasoning-intensive benchmar...
By Radin Hamidi Rad, Amin Bigdeli, Negar Arabzadeh, Sajad Ebrahimi, Charles L. A. Clarke, Benjamin C. M. Fung, Ebrahim Bagheri
arXiv:2607. 24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures.
By Seonok Kim
arXiv:2607. 22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space.
By Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang
Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.
arXiv:2606. 29648v1 Announce Type: cross Abstract: Different retrievers, including lexical, semantic, and multimodal approaches, provide highly complementary strengths for multimodal document understanding, yet most systems combine them through fixed pipelines that cannot adapt to the demands of individual reasoning steps.
By Bohan Yao, Shruthan Radhakrishna, Vikas Yadav
arXiv:2606. 05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference.
By Yongjie Wang, Xinyue Zhang, Kunhong Yao, Zhiwei Zeng, Kaisong Song, Jun Lin, Zhiqi Shen
arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.
By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.
By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu