The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.
By Mika Okamoto, Ansel Kaplan Erol
arXiv:2609.36228v1 Announce Type: new
Abstract: The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements...
By Zhen Tao, Alize Kahraman, Shidong Pan, Zhenchang Xing, Chiara Ullstein, Jens Grossklags, Chunyang Chen
arXiv:2607. 08292v1 Announce Type: cross Abstract: The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance toward machine-readable compliance artifacts.
By Lea Roxanne Muth, Marian Margraf
arXiv:2606. 32004v1 Announce Type: new Abstract: Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks.
By Sameer Malik, Ayush Singh, Amar Prakash Azad
Code-as-Auditor is an LLM-based framework that transforms regulatory information into formal checklists and executable decision trees, encoding rules as interpretable code. During inference, the model expands each checklist item into factual and counterfactual questions, guiding reasoning over case-specific evidence and potential violations. This pipeline moves from evidence identification to rule application and final decision-making, with a self‑verification loop that enhances logical consistency and traceability, leading to more accurate and evidence‑backed compliance evaluations in privacy and data protection scenarios.
By Jisoo Kim, Taeyoon Kwack, Jinwoo Jang, Woo Kyung Kim, Honguk Woo
arXiv:2607. 08288v1 Announce Type: cross Abstract: In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed for risk assessment and compliance.
By Lea Roxanne Muth, Marian Margraf