arXiv AI By Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Read the original on arXiv AI →

HyperBrowseComp is a multilingual and multimodal web‑browsing benchmark featuring 423 manually authored, human‑validated questions in 13 languages. The questions are intentionally difficult, requiring users to locate obscure evidence, follow multi‑step clue chains, or inspect heterogeneous sources such as videos, scanned documents, images, or maps. The benchmark filters out easier questions by testing models without internet access and evaluates performance using provider‑native search and a shared external retrieval harness, with a human evaluation on a sample to contextualize model effort.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks

WebArxiv is a reproducible benchmark designed to evaluate multimodal web agents on arXiv-related tasks. It consists of 510 static, time‑invariant tasks that require multi‑constraint paper retrieval, fine‑grained content extraction, and cross‑paper comparison, each with a deterministic ground truth. The benchmark highlights challenges for foundation‑model agents, such as over‑reliance on fixed interaction histories, and introduces a lightweight dynamic‑memory mechanism to improve adaptive retrieval and reasoning.

By Zihao Sun, Zijing Shi, Ling Chen
arXiv AI
Sep 1

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However...

By Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Hui Huang, Donghao Zhou, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Zhaoxiang Zhang, Qiangpeng Yang, Shilei Wen