arXiv AI

PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines

arXiv:2606. 32004v1 Announce Type: new Abstract: Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks.

Hugging Face Trending Papers
Aug 3

CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMs

Trust is fundamental in modern regulatory ecosystems, and compliance checking plays a critical role in fostering that trust. Regulatory compliance verification is essential for businesses operating in highly controlled environments, as it ensures alignment with sector-specific guidelines across domains such as financial reporting, data privacy, and cybersecurity.

arXiv AI
Aug 26

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

Granite.Trust Policy Tools introduces a YAML-based Actionable Policy schema that specifies what content a generative AI model can or cannot produce, allowing exception-based governance. It also offers a synthetic data generation pipeline to create policy-aligned training data and a suite of tools for defining and enforcing these policies throughout the AI lifecycle. The tools and example policies are open source, enabling organizations to tailor safety policies to their specific risks and regulatory contexts.

By Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox
arXiv AI
Sep 18

Code-as-Auditor: Executable Compliance Reasoning via Regulation-to-Code

Code-as-Auditor is an LLM-based framework that transforms regulatory information into formal checklists and executable decision trees, encoding rules as interpretable code. During inference, the model expands each checklist item into factual and counterfactual questions, guiding reasoning over case-specific evidence and potential violations. This pipeline moves from evidence identification to rule application and final decision-making, with a self‑verification loop that enhances logical consistency and traceability, leading to more accurate and evidence‑backed compliance evaluations in privacy and data protection scenarios.

By Jisoo Kim, Taeyoon Kwack, Jinwoo Jang, Woo Kyung Kim, Honguk Woo
arXiv AI
Jun 26

Autoformalization of Agent Instructions into Policy-as-Code

arXiv:2606. 26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned classifiers, prompt-based steering) that offer no formal guarantees, or on hand-coded symbolic enforcement that does not scale to the breadth of real policy specifications.

By Adam Mondl, Matthew Maisel, John H. Brock
arXiv Computation and Language
Aug 25

Grounded Normative Rule Generation with Structured Search

The paper introduces Grounded Normative Rule Generation (GNRS) and a new framework called GNRS-Search that uses Markov Chain Monte Carlo sampling to optimize a discrete And-Or Graph for rule synthesis. By separating operational feasibility from prose generation, the method localizes rule failures before final text creation. Evaluations on GNRS-Bench and RealCharter-Bench show significant improvements in rubric quality and executable metrics, demonstrating that the gains come from robust operational logic rather than stylistic tuning.

By Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao, Shuo Chen, Song-Chun Zhu
arXiv AI
Sep 17

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.

By Mika Okamoto, Ansel Kaplan Erol
Hugging Face Trending Papers
Aug 6

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency.