arXiv:2607. 14443v1 Announce Type: new Abstract: Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor layer: they look at screenshots, predict coordinates, click, and hope that the visible state changed as intended.
By Yong Liu, Zhenyi Zhong, Zhanpeng Shi
The paper introduces ASIL, an Agent‑Software Interaction Layer that replaces traditional screenshot‑and‑click interfaces with structured JSON observations and code‑executable semantic actions. ASIL is implemented across 15 applications and evaluated on 300 single‑application and 80 multi‑application tasks, achieving over 80% success with fewer than five actions per task. The structured interface also improves training efficiency, boosting performance of Qwen models from 58–66% to 72–80% with small‑scale supervised fine‑tuning and further gains with on‑policy reinforcement learning.
By Rui Xie, Lu Chen
arXiv:2608. 09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout.
By Santosh Patapati
The study evaluates computer-use agents (CUAs) for blind users by conducting a three‑week diary study with eight participants using the OLLA prototype. Across 1,258 commands in 12 desktop applications, GPT‑5 achieved the highest success rate of 52.5%, while analysis uncovered failures in grounding, planning, constraint‑tracking, and termination. Interviews highlighted additional needs beyond automation for blind users.
By Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
The paper introduces a safety‑bounded gateway that translates IEEE 11073 Service‑Oriented Device Connectivity (SDC) into the Model Context Protocol (MCP) for medical AI agents. It exposes device metrics, alarms, context references, and semantic metadata as read‑only resources, while representing selected action affordances as policy‑validated dry‑run tools, ensuring that agent requests never trigger actual device operations. A Python prototype demonstrates fault and lifecycle experiments, deterministic baselines, and multi‑model agent evaluation, showing improved semantic conformity and preservation of the no‑execution boundary.
By Bennet Gerlach, Stefan Fischer
arXiv:2605. 10555v2 Announce Type: replace Abstract: As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-oriented CRUD paradigms.
By Kai Pan, Rong Hou
arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
By Bo Zhang, Borui Zhang, Chenghao Jiang, Minglei Shi, Xiaofeng Wang, Zheng Zhu, Jie Zhou, Jiwen Lu
UI‑Venus‑2 is a general‑purpose foundation GUI agent that operates across mobile, web, and desktop environments using a unified closed‑loop reasoning‑action framework. The report details how the system expands environment coverage to over 170 multilingual mobile apps and native desktop OSes, scales task generation through a deep‑research pipeline, and enhances verification with trace‑level and sample‑level evaluators that use visual keypoints and multi‑model voting. Safety‑aware mechanisms are also incorporated to control consequential actions, positioning UI‑Venus‑2 as an efficient, open‑source tool for more generalizable, verifiable, and self‑reflective agents in real‑world applications.
By Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
arXiv:2609.24362v1 Announce Type: new
Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language m...
By Hexiong Yang, Mingrui Chen, Jie Cao, Ran He
arXiv:2607. 24770v1 Announce Type: new Abstract: Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical actions.
By Azizul Zahid, Subrata Biswas, Bashima Islam, Sai Swaminathan
arXiv:2608. 08939v1 Announce Type: new Abstract: The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices.
By Rahul Deivasigamani, Sayeda Faatin Alvi, Derqui Andrea, Kaushal Punjabi, Stjepan Picek
JarvisGUI is a new benchmark that tests GUI agents on cross-device workflows involving Android, Windows, and Ubuntu, requiring transfer of intermediate results and coordination across heterogeneous platforms. It formulates tasks as input-output transformations under a lightweight type system, enabling automatic composition of multi-step, cross-device workflows and dynamic evaluation within a unified framework. The benchmark reveals that state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a critical capability gap invisible to existing benchmarks.
By Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang, Zhiyin Lin, Zihan Li, Yuhang Guo, Yunhong Wang, Haifeng Wang