arXiv:2606. 29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observation interface in computer-use (CU) agents.
By Bojie Li, Noah Shi
arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.
By Ayae Ide, Tory Park, Jaron Mink, Tanusree Sharma
arXiv:2608. 09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout.
By Santosh Patapati
arXiv:2609.15696v1 Announce Type: cross
Abstract: Generative AI (GenAI) tools are increasingly woven into how blind and low-vision (BLV) people communicate, not only with digital information, but wit...
By Protik Dey, Mohd Saifuzzaman, Taslima Akter
Affora is a design system aimed at making software interfaces more readable by computer-use agents while still allowing designers visual freedom and maintaining familiar human workflows. The authors conducted three controlled studies on component implementations, visual variation, and interaction-design principles, using the results to create guidance from individual components to full sites, along with reusable implementations and executable checks. Evaluation on independently authored interfaces showed performance gains where Affora addressed existing deficits, with limited effects elsewhere, and a workflow case suggested reduced interaction cost.
By Jin Gao
arXiv:2606. 14777v1 Announce Type: cross Abstract: Many moments in the real world do not wait for a user to ask.
By Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Haowen Hou, Zheming Liang, Congcong Wang, Yuhang Cao, Shenglong Ye, Shuai Xie, Shuhuan Gu, Haoyang Huang, Qingyi Si, Nan Duan, Jiaqi Wang
arXiv:2504. 17331v3 Announce Type: replace-cross Abstract: Locomotion plays a crucial role in shaping the user experience within virtual reality environments.
By Suleyman Ozdel, Kadir Burak Buldu, Enkelejda Kasneci, Efe Bozkir
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
By Chen Chen, Zhehuai Chen
The study evaluates computer-use agents (CUAs) for blind users by conducting a three‑week diary study with eight participants using the OLLA prototype. Across 1,258 commands in 12 desktop applications, GPT‑5 achieved the highest success rate of 52.5%, while analysis uncovered failures in grounding, planning, constraint‑tracking, and termination. Interviews highlighted additional needs beyond automation for blind users.
By Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
The technical report introduces Gander, an end‑to‑end model that integrates omni perception, real‑time interaction, and agentic capabilities into a single framework. Unlike traditional turn‑based systems, Gander continuously processes streaming inputs from video, speech, and text, enabling natural full‑duplex interaction in both everyday conversations and workflow‑oriented scenarios. Its architecture features a Cerebellum‑Brain collaboration—where the Cerebellum handles real‑time interaction and omni conversational tasks while the Brain manages complex reasoning—and a streaming Thinker‑Talker design that flattens inputs and outputs into an ordered token stream for low‑latency, continuous dialogue. Evaluations across conversational ability, omni understanding, interactive capability, and agentic intelligence show that Gander matches state‑of‑the‑art open‑source models in spoken dialogue while maintaining robust performance in noisy, multi‑party, and backchannel environments.
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2608. 14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce.
By Chen Chen, Zhehuai Chen
AgentLens is a mobile GUI agent that adapts its visual communication during task execution, offering Full UI, Partial UI, and GenUI modalities. It uses a Virtual Display to allow background operation while selectively overlaying visual information. In a study with 21 participants, 85.7% preferred AgentLens, which also scored highest on usability and adoption intent.
By Jeonghyeon Kim, Byeongjun Joung, Junwon Lee, Joohyung Lee, Taehoon Min, Sunjae Lee