arXiv AI

Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer

The paper argues for a Foundation Model Operating System (FMOS) to virtualize foundation model interactions, similar to how operating systems abstract hardware. Current AI stacks are fragmented, with each framework embedding its own runtime for state, memory, budgets, and guardrails, leading to non-portable behavior and brittle governance. An FMOS would orchestrate knowledge across memory tiers, manage model selection and resource allocation, and enforce verification and policy, learning when to intervene or allow direct inference based on operational experience.

arXiv AI
Sep 25

AgentKernel: The Trust-Native Agentic Operating System

AgentKernel proposes a trust‑native operating system for AI agents, arguing that current governance layers are insufficient because they share the same process trust boundary as the agents. The OS introduces a mandatory enforcement boundary organized into four pillars—Identity, Perception, Cognition, and Execution—each adapting classical OS security principles to address semantic‑level failures such as prompt injection, memory poisoning, and tool misuse. By wrapping the agent lifecycle in this structured, non‑bypassable framework, AgentKernel aims to provide a unified security layer that can enforce identity, input mediation, memory governance, and execution control across the entire agent lifecycle.

By Zhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao, Shuo Li, Zhuotao Liu
arXiv AI
Jul 29

Towards an Agent Operating System - Lessons from Classical and Cloud OS

arXiv:2607. 25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, followed by the articulation of a small set of stable abstractions with well-defined semantics, and finally consolidation around those abstractions into a platform that applications can portably target.

By Gosia Steinder, Hubertus Franke
arXiv AI
Sep 11

A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model

The paper presents a taxonomy of architecture options for foundation-model-based agents, covering functional capabilities, non‑functional qualities, and operational aspects of design‑time and run‑time phases. It also introduces a decision model to guide critical design and runtime choices, aiming to streamline and improve the development of such agents. By unifying these classifications, the authors seek to reduce fragmentation in the field and provide a structured framework for architects and developers.

By Jingwen Zhou, Qinghua Lu, Jieshan Chen, Liming Zhu, Xiwei Xu, Zhenchang Xing, Stefan Harrer
arXiv AI
Jun 12

The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements

arXiv:2606. 12797v1 Announce Type: new Abstract: Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and financial advising.

By Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu, Nirwan Ansari
Hugging Face Trending Papers
Jun 2

Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents

Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, request human authority, generate tools, and perform side effects that must be resumed and audited. This paper presents Agent libOS, a library-OS-inspired runtime substrate for LLM agents.

arXiv AI
Jul 22

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance

arXiv:2607. 18246v1 Announce Type: new Abstract: We present Phionyx, a deterministic AI runtime architecture derived from the broader Echoism interaction framework that introduces a governance-first approach to AI engineering: treating large language model (LLM) outputs as noisy sensor measurements rather than direct decisions.

By Ali Toygar Abak