The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.
By Mika Okamoto, Ansel Kaplan Erol
arXiv:2609.36228v1 Announce Type: new
Abstract: The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements...
By Zhen Tao, Alize Kahraman, Shidong Pan, Zhenchang Xing, Chiara Ullstein, Jens Grossklags, Chunyang Chen
arXiv:2607. 08292v1 Announce Type: cross Abstract: The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance toward machine-readable compliance artifacts.
By Lea Roxanne Muth, Marian Margraf
arXiv:2606. 32004v1 Announce Type: new Abstract: Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks.
By Sameer Malik, Ayush Singh, Amar Prakash Azad
Code-as-Auditor is an LLM-based framework that transforms regulatory information into formal checklists and executable decision trees, encoding rules as interpretable code. During inference, the model expands each checklist item into factual and counterfactual questions, guiding reasoning over case-specific evidence and potential violations. This pipeline moves from evidence identification to rule application and final decision-making, with a self‑verification loop that enhances logical consistency and traceability, leading to more accurate and evidence‑backed compliance evaluations in privacy and data protection scenarios.
By Jisoo Kim, Taeyoon Kwack, Jinwoo Jang, Woo Kyung Kim, Honguk Woo
arXiv:2607. 08288v1 Announce Type: cross Abstract: In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed for risk assessment and compliance.
By Lea Roxanne Muth, Marian Margraf
arXiv:2601. 22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes.
By Daniel Commey
arXiv:2605.29742v2 Announce Type: replace
Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered auth...
By Yeong-Joon Ju, Seong-Whan Lee
The paper examines how large language models (LLMs) can assist in creating regulatory compliance artifacts for EU sustainability and privacy laws, specifically Digital Product Passports (DPPs) under the Ecodesign for Sustainable Products Regulation and Data Protection Impact Assessments (DPIAs) under the General Data Protection Regulation. It investigates the effects of data extraction instructions and regulatory ambiguity on the quality and consistency of LLM-generated artifacts, benchmarking various models against manually crafted ground‑truth schemas. Findings indicate that looser guidelines, like those for DPIAs, demand more extensive prompts to achieve consistency, whereas stricter formatting rules for DPPs yield consistent outputs but may introduce hallucinations.
By Adriana Watson, Marco B\"ucheler, Grant Richards
arXiv:2608. 07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls.
By Allison Wilson, Sina Moradi Sabet, Diar Shakimov, Panteha Shahrivar, Mohammad Reza Bagheri, Dean Konenkamp, Mohammad A. Tayebi
arXiv:2606. 02965v2 Announce Type: replace Abstract: As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2608. 20204v1 Announce Type: new Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs.
By Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean