The paper discusses how financial institutions are increasingly using AI agents in areas such as credit, fraud, and compliance, yet current governance focuses only on individual components. It introduces ARIA, a finance‑specific reference architecture that adds six capabilities—policy specification, population‑level monitoring, bounded authority, runtime containment, adaptive policy change, and preserved human oversight—to address the gap of constitutional non‑compositionality. Two simulations demonstrate how local controls can miss collective bias and how observed‑versus‑expected monitoring can provide earlier warnings of drift.
By Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Diaz, Pablo Delgado Romero
arXiv:2608. 16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task.
By Bowen Li, Guojun Wang
Financial institutions are beginning to deploy agentic workflows in credit, fraud, collections, compliance, and operational control. Governance remains largely component-centric: each model or agent i...
The paper "Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance" presents a taxonomy of twenty inference‑time mechanisms for monitoring, verification, and enforcement, each evaluated on a four‑point readiness scale using evidence from four vendors. It applies this taxonomy to a two‑dimensional adversary model and maps the mechanisms to four governance scenarios, finding that most mechanisms are commercially available but only adequate against cooperative or low‑to‑medium‑capability users, not high‑capability state‑level deployers. The study also links inference‑stage controls to hardware‑stage mechanisms through a substitution principle and reports a second‑rater reliability of 0.74.
whyItMatters":"The work identifies the current gaps and readiness of inference‑time governance tools, highlighting that existing mechanisms are insufficient against powerful adversaries and thus informing future regulatory and technical development."
By Samar Ansari
The paper "When Agents Act Unwatched: The Reduced‑Supervision Paradox in Agentic AI" discusses how the promise that AI systems will continue acting after users stop watching creates an accountability inversion. It argues that as stepwise supervision recedes, verification shifts into the runtime infrastructure—authority, records, interrupts, outcome checks, and repair—forming what the authors call the reduced‑supervision paradox. A 63‑artifact audit across research papers and engineering sources shows that agents’ action surfaces are more visible than the mechanisms needed to hold them accountable, with tool mediation and monitoring traces appearing in 40 and 37 artifacts, while checkpoint placement, validator independence, recovery, and contestability are rarely visible.
"whyItMatters":"The study highlights that observable action paths can replace accountability when verification is moved onto users after meaningful intervention is no longer possible."
By Hanjing Shi, Dominic DiFranzo
Compute governance today is a governance of training: the thresholds, reporting requirements, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulat...
arXiv:2609.37457v1 Announce Type: new
Abstract: Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate actio...
By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
The paper introduces a compliance-first AI architecture for regulated finance, treating regulation as an orientation layer rather than a rigid rule set. It uses a regulatory intent matrix and a governed policy compiler to translate regulatory requirements into concrete prohibitions, obligations, and runtime budgets, while maintaining proportional committee activation and supervisory oversight. Evidence and decisions are recorded on a permissioned DAG with deterministic timestamps, enabling replay, provenance checks, and clear attribution of failures, and clause-level legal indexing ensures portability across the DACH region and the EU.
By Walter Kurz, Reinhard Magg
arXiv:2606. 30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows.
By Anuj Kaul, Qianlong Lan, Pranay Gupta
arXiv:2606. 30970v2 Announce Type: replace Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows.
By Anuj Kaul, Qianlong Lan, Pranay Gupta
The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.
By Ali Atiah Alzahrani
arXiv:2608. 16402v1 Announce Type: new Abstract: Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a goal.
By Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar, Bhavesh Gadhe