arXiv:2608.22538v1 Announce Type: new
Abstract: Policy-governed agents must interpret case evidence while following an authorized procedure. We present \textsc{Stage}, an executable-graph framework t...
By Mengxi Luo, Changjia Chen, An Cao, Zirong Huang, Wanyi Dai
The paper introduces a governed approach to enterprise analytics in which a language model interprets user queries and a deterministic policy selects and runs pre‑approved analytical programs that return both results and evidence. The authors demonstrate that this restriction remains expressive for a defined analytical class—including relational operations, aggregation, comparison, windows, ranking, and similarity—while ensuring reproducibility through fixed meaning, policy, data, and execution rules. In experiments with 440 runs, three 8B models generated SQL and selected tools at runtime, whereas a policy‑executed analyzer achieved a perfect 110/110 match across all test datasets, though no runtime‑planning episodes matched the full answer‑and‑evidence contract.
"whyItMatters":"The study shows that a governed, policy‑driven framework can reliably produce accurate, reproducible analytics results, highlighting a viable path for controlled AI‑driven data analysis."
By MasterControl AI Lab
arXiv:2609.14400v1 Announce Type: new
Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assum...
By Hongliu Cao
arXiv:2609.37457v1 Announce Type: new
Abstract: Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate actio...
By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
arXiv:2608.29971v1 Announce Type: new
Abstract: As agentic systems evolve into complex multi agent orchestration workflows, there is a growing and critical need for systematic frameworks that measure...
By Ram Kulathumani, Regunathan Radhakrishnan, Anupam Tripathi, Xiangbo Mao, Roshanak Omrani, Keshav Somani, Shwet Kamal Mishra, Shayna Lurya
The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.
By Adam Mazzocchetti
The paper introduces Graph‑Agentic Retrieval‑Augmented Generation (RAG), a system that blends structured evidence with adaptive agents capable of planning retrieval, navigating relations, verifying claims, delegating tasks, and employing tools. It highlights how defects in graph construction can propagate through retrieval and control decisions, potentially leading to significant outcomes. To address these risks, the authors propose an assurance‑by‑construction framework with five interface contracts—evidence, retrieval, reasoning, capability & delegation, and outcome—that make provenance, validity, authorization, uncertainty, and recoverability explicit, and outline an evaluation agenda for social‑good applications.
By Vijay Bommireddy, Raviteja Bommireddy
arXiv:2609.37454v1 Announce Type: new
Abstract: Underwriters in commercial Property and Casualty (P&C) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a...
By Vivek Kumar Singh, Gautam Bhowmick
arXiv:2607. 25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.
By Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen
arXiv:2607. 09175v1 Announce Type: new Abstract: Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness.
By Dan C. Hsu, Luke Lu
arXiv:2609.32754v2 Announce Type: replace
Abstract: Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long seq...
By Jiecong Wang, Hao Peng, Zhanyi Wang
arXiv:2606. 26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications.
By Adam Mondl, Matthew Maisel, John H. Brock