arXiv AI

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

arXiv:2608. 03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools.

arXiv AI
Sep 25

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.

By Ivan Matveev
arXiv AI
Aug 28

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

The paper introduces ASIL, an Agent‑Software Interaction Layer that replaces traditional screenshot‑and‑click interfaces with structured JSON observations and code‑executable semantic actions. ASIL is implemented across 15 applications and evaluated on 300 single‑application and 80 multi‑application tasks, achieving over 80% success with fewer than five actions per task. The structured interface also improves training efficiency, boosting performance of Qwen models from 58–66% to 72–80% with small‑scale supervised fine‑tuning and further gains with on‑policy reinforcement learning.

By Rui Xie, Lu Chen
arXiv AI
Sep 25

SheetMind: Actions Set Accuracy, Agents Set the Failure Mode

SheetMind is a Manager‑Action‑Reflection framework that evaluates how much spreadsheet agent performance derives from the agents themselves versus the shared action interface. In a controlled study on all 221 tasks of the SheetCopilot Benchmark, replacing the high‑level action API with primitive cell operations drops accuracy by 47.1 points, while adding a Reflection Agent improves performance by 4.5 points and a Manager by 1.4 points. The framework also shows that decomposition changes failure modes, reducing silent wrong outputs from 33% to 25%, and that GPT‑5 and GPT‑5‑mini achieve similar performance, whereas GPT‑3.5 underperforms significantly.

By Lyuhao Chen, Xi Cheng, Yanming Kang, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Yulang Fei, Brian Zhu, Daniel Jin, Binze Cai, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao