arXiv AI

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

arXiv:2606. 20724v2 Announce Type: replace Abstract: Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, and terminate confidently while still missing fields, over-including unsupported items, or relying on stale evidence.

arXiv Computation and Language
Sep 11

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

OpenResearcher is a fully open, reproducible pipeline for generating long‑horizon deep research trajectories that interleave search, evidence aggregation, and multi‑step reasoning. It decouples corpus bootstrapping from trajectory synthesis and runs the search‑and‑browse loop offline using three browser primitives over a 15M‑document corpus. Using GPT‑OSS‑120B as a teacher, the pipeline produced over 97K trajectories, enabling a 30B‑A3B model to achieve 54.8% accuracy on BrowseComp‑Plus and providing insights into pipeline design through controlled analysis.

By Zhuofeng Li, Dongfu Jiang, Xueguang Ma, Haoxiang Zhang, Ping Nie, Yuyu Zhang, Kai Zou, Jianwen Xie, Yu Zhang, Wenhu Chen
arXiv AI
Aug 19

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.

By AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou
arXiv AI
Sep 25

WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks

WebArxiv is a reproducible benchmark designed to evaluate multimodal web agents on arXiv-related tasks. It consists of 510 static, time‑invariant tasks that require multi‑constraint paper retrieval, fine‑grained content extraction, and cross‑paper comparison, each with a deterministic ground truth. The benchmark highlights challenges for foundation‑model agents, such as over‑reliance on fixed interaction histories, and introduces a lightweight dynamic‑memory mechanism to improve adaptive retrieval and reasoning.

By Zihao Sun, Zijing Shi, Ling Chen
arXiv AI
Aug 11

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.

By Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira, Atsuki Sato, Tatsumi Sunada, Shota Onohara, Hiromasa Yamanishi, Mashiro Toyooka, Kunato Nishina, Ryoma Maeda, Kiyoharu Aizawa, Toshihiko Yamasaki
arXiv Computation and Language
Sep 24

Improving LLM-based Autonomous Web Agents with Filtering

The paper investigates how to improve large language model (LLM) based autonomous web agents by filtering irrelevant webpage content. The authors reproduce baseline models on the WebArena benchmark and identify failure modes caused by raw HTML input. They propose DeBERTa‑ and T5‑based retrieval models that rank HTML elements by relevance, fine‑tuned on Mind2Web data, and demonstrate that the DeBERTa model raises the LLaMA‑2‑70B agent’s success rate from 1.97% to 2.96%. Additionally, a zero‑shot ColBERT retriever achieves recall of 0.52 on Mind2Web and 0.47 on WebArena.

By Zhitong Guo, Jing Yu Koh, Ruiyu Li
arXiv AI
Aug 7

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

arXiv:2608. 06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging.

By Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
arXiv AI
Aug 7

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv:2608. 05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers.

By Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
arXiv AI
3d ago

TRACE: Trajectory Selection for Parallel Scaling of Search Agents

TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.

By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan
arXiv Machine Learning
Aug 12

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

arXiv:2607. 24850v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons.

By Lang Mei, Xiaohan Yu, Chong Chen, Liyan Liu, Xiangnan Chen, Jinchao Ma, Chao Feng, Li Huang, Siyu Mo, Sichen Kang, Yunkun Xu, Zhihan Yang, Zhujun Xue, Jingren Zhang, Qing He, Yingdi Huang, Hao Jiang, Ziao Ma, Zewei Pan, Minhao Sun, Zhuo Tao, Jinzhao Xiao, Gangtao Xin, Huanyao Zhang, Wenjian Zhang, Jiangshan Zhang, Guojie Zhu, Fangzhou Zou, Jiaxin Mao, Wentao Zhang