Hugging Face Trending Papers

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

arXiv AI
Sep 4

MasterControl Seventeen Every Time

The paper introduces a governed approach to enterprise analytics in which a language model interprets user queries and a deterministic policy selects and runs pre‑approved analytical programs that return both results and evidence. The authors demonstrate that this restriction remains expressive for a defined analytical class—including relational operations, aggregation, comparison, windows, ranking, and similarity—while ensuring reproducibility through fixed meaning, policy, data, and execution rules. In experiments with 440 runs, three 8B models generated SQL and selected tools at runtime, whereas a policy‑executed analyzer achieved a perfect 110/110 match across all test datasets, though no runtime‑planning episodes matched the full answer‑and‑evidence contract. "whyItMatters":"The study shows that a governed, policy‑driven framework can reliably produce accurate, reproducible analytics results, highlighting a viable path for controlled AI‑driven data analysis."

By MasterControl AI Lab
arXiv AI
Sep 1

EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL Configuration

arXiv:2608.29971v1 Announce Type: new Abstract: As agentic systems evolve into complex multi agent orchestration workflows, there is a growing and critical need for systematic frameworks that measure...

By Ram Kulathumani, Regunathan Radhakrishnan, Anupam Tripathi, Xiangbo Mao, Roshanak Omrani, Keshav Somani, Shwet Kamal Mishra, Shayna Lurya
arXiv AI
Aug 19

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.

By Adam Mazzocchetti
arXiv AI
Sep 10

Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction

The paper introduces Graph‑Agentic Retrieval‑Augmented Generation (RAG), a system that blends structured evidence with adaptive agents capable of planning retrieval, navigating relations, verifying claims, delegating tasks, and employing tools. It highlights how defects in graph construction can propagate through retrieval and control decisions, potentially leading to significant outcomes. To address these risks, the authors propose an assurance‑by‑construction framework with five interface contracts—evidence, retrieval, reasoning, capability & delegation, and outcome—that make provenance, validity, authorization, uncertainty, and recoverability explicit, and outline an evaluation agenda for social‑good applications.

By Vijay Bommireddy, Raviteja Bommireddy
arXiv AI
Jun 26

Autoformalization of Agent Instructions into Policy-as-Code

arXiv:2606. 26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications.

By Adam Mondl, Matthew Maisel, John H. Brock