arXiv:2608. 08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers.
By Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang
arXiv:2603.04949v2 Announce Type: replace
Abstract: As web agents close the gap with humans on benchmarks, one question arises: Do today's agents perform just as well on tomorrow's web? We introduce...
By Md Farhan Ishmam, Kenneth Marino
arXiv:2604. 08523v2 Announce Type: replace-cross Abstract: AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites?
By Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Zhuofeng Li, Xingwei Qu, Zhengkang Guo, Yuanzhe Shen, Dingjie Song, Han Zhou, Tuney Zheng, Xian Wu, Hao Yu, Songcheng Cai, Yi Lu, Yunzhuo Hao, Minyi Lei, Liang Chen, Kai Zou, Huifeng Yin, Wendong Xu, Dongfu Jiang, Ping Nie, Jiaheng Liu, Wenhu Chen, Kelsey R. Allen
arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.
By Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira, Atsuki Sato, Tatsumi Sunada, Shota Onohara, Hiromasa Yamanishi, Mashiro Toyooka, Kunato Nishina, Ryoma Maeda, Kiyoharu Aizawa, Toshihiko Yamasaki
arXiv:2606. 09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.
By Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan
arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
By Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov
The paper investigates how large‑language‑model (LLM) based AI agents mix latency, local resource usage, and container bottlenecks when processing user requests that involve remote LLM calls and local tool execution. By measuring three representative tasks—retrieval‑augmented question answering, web search, and software coding—the authors show that agents exhibit diverse resource dynamics, with concurrent requests revealing task‑specific bottlenecks in CPU, disk I/O, and memory. Leveraging these insights, they propose CPU‑aware tool admission and task‑aware CPU allocation, achieving up to a 5.4× speed‑up for CPU‑sensitive tasks and a 32% reduction in average latency across multiple tasks.
By Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang
arXiv:2609.35814v1 Announce Type: cross
Abstract: As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. Thi...
By Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao, Daisy Xinlei Lin, Royce Cheng-Yue, Keagan Long, Kyle Wong, Bhuwan Dhingra, Xiangjun Wang, Shuyan Zhou
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.
arXiv:2606. 01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents.
By Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
By Chen Chen, Zhehuai Chen
The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.
By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata