arXiv AI

PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs

arXiv:2608. 09028v1 Announce Type: new Abstract: Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints.

arXiv AI
Jun 9

From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing

arXiv:2606. 08932v1 Announce Type: cross Abstract: Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but silently drops nested exceptions or counter-exceptions, producing outputs that appear compliant yet break on important edge cases.

By Jian Chen, Siyuan Li, Chucheng Wan, Zixuan Yuan
arXiv AI
Sep 12

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is a benchmark that evaluates how well large language models (LLMs) understand and apply version-constraint resolution semantics, such as determining whether a version satisfies constraints like ^1.2.3 or >=2.0. The study finds that many models struggle with certain corner cases, with GPT‑5.1 performing poorly while Claude and Opus perform much better. The authors suggest that the failures stem from an activation/application gap rather than a lack of knowledge, and recommend that coding agents delegate version resolution to a dedicated resolver tool.

By Qibai Chen, Zeming Liu
arXiv AI
Jul 10

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

arXiv:2607. 08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic Process Analysis (STPA).

By Samuel Tetteh, Udip Shrestha, Joshua R. Waite, Cody Fleming
arXiv Machine Learning
Sep 18

A Policy Profile for Croissant: Refusal as a Property of the Dataset

Croissant, a JSON‑LD based descriptor for machine‑learning datasets, now includes a policy profile that specifies how data‑use conditions are evaluated. The profile defines five operators with explicit decision procedures, enabling a gate to determine admissible operations from the descriptor alone and record what was checked. Experiments on two corpora and a real nf‑core pipeline show that the profile’s decisions match native descriptors, add minimal overhead, and cover all operators, refusal classes, and conformance clauses.

By Alexander Chernov