Data Flow Control: Data Safety Policies for AI Agents
arXiv:2606. 05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users.
Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness, correctness is not safety.
arXiv:2606. 05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users.
arXiv:2606. 08661v1 Announce Type: cross Abstract: Data agents integrate LLM-driven reasoning with relational data access, executable analytical tools, and multi-step workflow orchestration, making them increasingly central to enterprise analytics.
arXiv:2605. 21027v2 Announce Type: replace-cross Abstract: Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional business intelligence tools or Text-to-SQL systems.
arXiv:2609.35807v1 Announce Type: cross Abstract: LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/...
arXiv:2606. 26627v1 Announce Type: cross Abstract: Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf.
arXiv:2604. 14401v2 Announce Type: replace Abstract: Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving conditions.
The paper introduces two zero‑trust frameworks for cloud data engineering and analytical processing. The first, Zero‑Trust Agentic Data Engineering, automatically generates, deploys, and verifies complete data‑engineering solutions from natural‑language tasks, requiring evidence from repositories, deployments, runtimes, and policies. The second, Zero‑Trust Agentic OLAP, combines governed data preparation with verified online analytical processing, allowing production promotion only after rigorous validation and evidence‑bound approval, and ensuring analytical outputs are released only after same‑snapshot execution, exact result equivalence, deterministic grounding, and reflection. Both frameworks rely on three core abstractions—graph engineering for evidence‑gated workflow structure, loop engineering for bounded recovery, and agent‑harness engineering for zero‑trust execution—and are evaluated under nominal execution, controlled failures, bounded recovery, and policy‑constrained conditions to measure verified completion, recovery, authorization enforcement, production promotion, and verified OLAP execution.
arXiv:2508.15526v2 Announce Type: replace Abstract: The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabili...
arXiv:2608. 13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation.
arXiv:2609.14003v1 Announce Type: cross Abstract: Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provid...
Granite.Trust Policy Tools introduces a YAML-based Actionable Policy schema that specifies what content a generative AI model can or cannot produce, allowing exception-based governance. It also offers a synthetic data generation pipeline to create policy-aligned training data and a suite of tools for defining and enforcing these policies throughout the AI lifecycle. The tools and example policies are open source, enabling organizations to tailor safety policies to their specific risks and regulatory contexts.
arXiv:2606. 19319v1 Announce Type: cross Abstract: Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data.