arXiv AI By Yogeswar Reddy Thota

LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents

Read the original on arXiv AI →

arXiv:2606. 30697v1 Announce Type: cross Abstract: Current operating systems expose interfaces optimized for human users but not for AI agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

The paper introduces ASIL, an Agent‑Software Interaction Layer that replaces traditional screenshot‑and‑click interfaces with structured JSON observations and code‑executable semantic actions. ASIL is implemented across 15 applications and evaluated on 300 single‑application and 80 multi‑application tasks, achieving over 80% success with fewer than five actions per task. The structured interface also improves training efficiency, boosting performance of Qwen models from 58–66% to 72–80% with small‑scale supervised fine‑tuning and further gains with on‑policy reinforcement learning.

By Rui Xie, Lu Chen
arXiv AI
Sep 2

Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications

The study evaluates computer-use agents (CUAs) for blind users by conducting a three‑week diary study with eight participants using the OLLA prototype. Across 1,258 commands in 12 desktop applications, GPT‑5 achieved the highest success rate of 52.5%, while analysis uncovered failures in grounding, planning, constraint‑tracking, and termination. Interviews highlighted additional needs beyond automation for blind users.

By Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
arXiv AI
6d ago

A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents

The paper introduces a safety‑bounded gateway that translates IEEE 11073 Service‑Oriented Device Connectivity (SDC) into the Model Context Protocol (MCP) for medical AI agents. It exposes device metrics, alarms, context references, and semantic metadata as read‑only resources, while representing selected action affordances as policy‑validated dry‑run tools, ensuring that agent requests never trigger actual device operations. A Python prototype demonstrates fault and lifecycle experiments, deterministic baselines, and multi‑model agent evaluation, showing improved semantic conformity and preservation of the no‑execution boundary.

By Bennet Gerlach, Stefan Fischer