Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents
Read the original on arXiv Computation and Language →Online web agents frequently add memory, workflow, or skill modules to a base actor, which can boost performance but also consume test‑time tokens—a cost rarely reported. This study evaluates such augmentation under a fixed inference budget, comparing AWM, ASI, and ReasoningBank to a token‑matched vanilla baseline across four WebArena domains and three models (Gemini 3 Flash, GPT‑5.4‑mini, Qwen 3.6‑27B). The vanilla baseline consistently matches or outperforms the augmentation methods in overall success rate while often using fewer tokens, a trend also seen on WorkArena‑L1 with Qwen 3.6‑27B. The results suggest that skills and workflow memory may only be beneficial in specific domains, and that run‑to‑run variance should be reported as a core evaluation criterion for online web agents.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.