arXiv Computation and Language By Sina Hajimiri, Masih Aminbeidokhti, Jose Dolz, Ismail Ben Ayed, Issam H. Laradji, Spandana Gella, Nicolas Gontier

Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents

Read the original on arXiv Computation and Language →

Online web agents frequently add memory, workflow, or skill modules to a base actor, which can boost performance but also consume test‑time tokens—a cost rarely reported. This study evaluates such augmentation under a fixed inference budget, comparing AWM, ASI, and ReasoningBank to a token‑matched vanilla baseline across four WebArena domains and three models (Gemini 3 Flash, GPT‑5.4‑mini, Qwen 3.6‑27B). The vanilla baseline consistently matches or outperforms the augmentation methods in overall success rate while often using fewer tokens, a trend also seen on WorkArena‑L1 with Qwen 3.6‑27B. The results suggest that skills and workflow memory may only be beneficial in specific domains, and that run‑to‑run variance should be reported as a core evaluation criterion for online web agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 11

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.

By Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira, Atsuki Sato, Tatsumi Sunada, Shota Onohara, Hiromasa Yamanishi, Mashiro Toyooka, Kunato Nishina, Ryoma Maeda, Kiyoharu Aizawa, Toshihiko Yamasaki
arXiv AI
Jun 9

Skill Retrieval Augmentation for Agentic AI

arXiv:2604. 24594v3 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their native parametric capabilities.

By Weihang Su, Jianming Long, Qingyao Ai, Qiaozhi He, Yichen Tang, Changyue Wang, Yiteng Tu, Yingbo Wang, Yiqun Liu