arXiv:2511. 12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains.
By Genglin Liu, Shijie Geng, Sha Li, Hejie Cui, Sarah Zhang, Xin Liu, Tianyi Liu
arXiv:2608. 08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers.
By Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang
Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.
By AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou
arXiv:2603.04949v2 Announce Type: replace
Abstract: As web agents close the gap with humans on benchmarks, one question arises: Do today's agents perform just as well on tomorrow's web? We introduce...
By Md Farhan Ishmam, Kenneth Marino
arXiv:2604. 06367v2 Announce Type: replace-cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries.
By Guruprasad Viswanathan Ramesh, Asmit Nayak, Basieem Siddique, Kassem Fawaz
arXiv:2606. 17645v1 Announce Type: new Abstract: Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits one structured tool action.
By Shiqi He, Yue Cui, Feijie Wu, Xinyu Ma, Jiaheng Lu, Yaliang Li, Bolin Ding, Mosharaf Chowdhury
Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction...
arXiv:2510. 19838v2 Announce Type: replace Abstract: Autonomous web agents powered by large language models (LLMs) show strong potential for performing goal-oriented tasks such as information retrieval, report generation, and online transactions.
By Shiqi He, Yue Cui, Xinyu Ma, Yaliang Li, Bolin Ding, Mosharaf Chowdhury
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
By Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c
MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.
By Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland
arXiv:2609.35814v1 Announce Type: cross
Abstract: As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. Thi...
By Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao, Daisy Xinlei Lin, Royce Cheng-Yue, Keagan Long, Kyle Wong, Bhuwan Dhingra, Xiangjun Wang, Shuyan Zhou
arXiv:2508. 04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incrementally solve the given task.
By Thassilo M. Schiepanski, Nicholas Pi\"el