arXiv:2606. 17929v1 Announce Type: new Abstract: Computer-using agents drive real software through the screen -- clicking and typing -- but they solve every task from scratch: asked to repeat a task, an agent re-reads the screen, re-reasons every tap, and pays the full cost again.
By Bojie Li
arXiv:2607. 09493v1 Announce Type: new Abstract: Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive.
By Sanjana Pedada, Aditya Dhavala, Neelraj Patil
arXiv:2608. 16381v1 Announce Type: new Abstract: Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages.
By Zhenhang Nie (iFLYTEK Co., Ltd., Hefei, China), Gui Zheng (iFLYTEK Co., Ltd., Hefei, China), Xudong Sun (iFLYTEK Co., Ltd., Hefei, China), Tailong Zhu (iFLYTEK Co., Ltd., Hefei, China), Bin Zhang (iFLYTEK Co., Ltd., Hefei, China)
arXiv:2606. 24551v1 Announce Type: new Abstract: Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and permitted actions.
By Xiao Zhou, Siyue Zhang, Yilun Zhao, Jinbiao Wei, Tingyu Song, Arman Cohan, Chen Zhao
The paper surveys efficient GUI agents, emphasizing that practical deployment requires more than task success—it must also minimize context, computation, action budget, and runtime overhead. It reviews key efficiency dimensions—observation, memory, action, and planner/system—and identifies recurring strategies such as selective reading, global-to-local visual allocation, recoverable memory, verification-aware control, and hybrid runtimes. The authors highlight open challenges, including accurate verifier cost accounting, benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.
By Bizhe Bai, Jiakang Yuan, Hongming Wu, Xinyue Wang, Jie Ren, Siyao Chen, Yuchen Ya, Fan Bai, Pai Peng, Huafeng Qin, Tao Chen
The paper introduces environment‑probing curation, a deployment‑compatible method that equips asynchronous curator agents with read‑only world tools to verify, scope, and refresh candidate memories without retraining models. In a GitHub Copilot‑based harness, this approach improves pass rates on CLBench from 39% to 73%, boosts reward metrics, and reduces both query counts and task‑agent costs. Across six APEX management‑consulting tasks, the method consistently outperforms baselines, yielding higher rewards and fewer tool calls while maintaining a compact task‑time interface.
By Susheel Suresh, Hazel Mak, Sahil Bhatnagar, Chhaya Methani, Alejandro Gutierrez Munoz