arXiv AI By Quilee Simeon, Justin M. Wei, Yile Fan

Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts

Read the original on arXiv AI →

Octopus Protocol is a hardware onboarding framework that allows an AI coding agent to automatically discover, identify, and integrate hardware devices into an AI system. Using a single bootstrap command, the agent runs a five-stage pipeline to enumerate visible hardware, infer device capabilities, generate typed Model Context Protocol tools, produce the necessary code, and activate a live endpoint. The system maintains a persistent daemon that repairs deployment failures, enabling consistent, platform‑agnostic interfaces across diverse hosts without manual integration code.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

The paper investigates whether locally deployed large language models can automate hardware design workflows that involve repetitive, dependency-ordered operations using specialized tools. A Model Context Protocol (MCP) server is created to emulate a proprietary hardware design tool, and a benchmark tests single and multi-step edits, invalid requests, misspelled prompts, and multi-server contexts. Seven open-source models are evaluated across different pipeline choices, revealing that strong models can nearly fully cover expected calls, but reliability hinges on task structure and agent configuration, with comprehensive tool descriptions reducing failures and multi-agent setups aiding weaker models at the cost of extra calls.

By Leonardo Liparulo, Francesco Pierri
arXiv AI
6d ago

A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents

The paper introduces a safety‑bounded gateway that translates IEEE 11073 Service‑Oriented Device Connectivity (SDC) into the Model Context Protocol (MCP) for medical AI agents. It exposes device metrics, alarms, context references, and semantic metadata as read‑only resources, while representing selected action affordances as policy‑validated dry‑run tools, ensuring that agent requests never trigger actual device operations. A Python prototype demonstrates fault and lifecycle experiments, deterministic baselines, and multi‑model agent evaluation, showing improved semantic conformity and preservation of the no‑execution boundary.

By Bennet Gerlach, Stefan Fischer
arXiv AI
Jul 29

Towards an Agent Operating System - Lessons from Classical and Cloud OS

arXiv:2607. 25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, followed by the articulation of a small set of stable abstractions with well-defined semantics, and finally consolidation around those abstractions into a platform that applications can portably target.

By Gosia Steinder, Hubertus Franke
arXiv AI
Aug 6

Terminal Agents Suffice for Enterprise Automation

arXiv:2604. 00073v3 Announce Type: replace-cross Abstract: There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomously.

By Patrice Bechard, Orlando Marquez Ayala, Emily Chen, Jordan Skelton, Sagar Davasam, Srinivas Sunkara, Vikas Yadav, Sai Rajeswar