arXiv:2511. 07332v2 Announce Type: replace-cross Abstract: Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements.
By Aarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han L\`u, Johan Obando-Ceron, Juan A. Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, Sai Rajeswar
arXiv:2605. 13527v3 Announce Type: replace Abstract: Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines.
By Kangning Zhang, Shuai Shao, Qingyao Li, Jianghao Lin, Lingyue Fu, Shijian Wang, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
arXiv:2606. 29538v1 Announce Type: cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge.
By Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress. Raw interaction traces preserve such information but are long and noisy to condition on, whereas text-only skills often omit the visual state that makes a procedure applicable.
EgoArgus is a new, human‑annotated dataset that tests visual‑language models (VLMs) as situational assistants in five everyday dialogue‑video scenarios. It evaluates how well VLMs understand and decide when to intervene, especially when visual and textual cues are helpful, irrelevant, or conflicting. The study finds that current VLMs still struggle to reliably act as egocentric assistants and that existing modality‑bias mitigation methods offer limited improvement.
By Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen
arXiv:2606. 11210v1 Announce Type: cross Abstract: Model Construction is a foundational practice in science learning that relies on visualization and interactivity.
By John Kos, Rudra Singh, Ashok Goel