Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.
By AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou
arXiv:2608. 08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers.
By Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang
arXiv:2604. 06367v2 Announce Type: replace-cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries.
By Guruprasad Viswanathan Ramesh, Asmit Nayak, Basieem Siddique, Kassem Fawaz
arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.
By Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira, Atsuki Sato, Tatsumi Sunada, Shota Onohara, Hiromasa Yamanishi, Mashiro Toyooka, Kunato Nishina, Ryoma Maeda, Kiyoharu Aizawa, Toshihiko Yamasaki
arXiv:2606. 15673v1 Announce Type: new Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement.
By Jiwan Chung, JiHyuk Byun, Vibhav Vineet, Seon Joo Kim
arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
By Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov
Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction...
arXiv:2511. 12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains.
By Genglin Liu, Shijie Geng, Sha Li, Hejie Cui, Sarah Zhang, Xin Liu, Tianyi Liu
WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.
By Guanqun Yang, Wei Yang, Xueqing Liu
arXiv:2603.04949v2 Announce Type: replace
Abstract: As web agents close the gap with humans on benchmarks, one question arises: Do today's agents perform just as well on tomorrow's web? We introduce...
By Md Farhan Ishmam, Kenneth Marino
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
By Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c
arXiv:2606. 05342v1 Announce Type: new Abstract: AI agents are increasingly asked to carry out work that spans minutes, hours, or longer.
By Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin, Hussein Mozzanar, Gagan Bansal, Maya Murad, Rafah Hosn, Saleema Amershi