arXiv AI

From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing

arXiv:2606. 08932v1 Announce Type: cross Abstract: Rule-following agents tasked with executing policies and regulations often fail via Silent Scope Omission (SSO): a model applies a general rule but silently drops nested exceptions or counter-exceptions, producing outputs that appear compliant yet break on important edge cases.

arXiv AI
Sep 17

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

The paper introduces the concept of Compositional Policy Violations (CPVs), where each step in an agentic AI workflow passes its individual compliance check, yet the overall execution violates higher‑level policies such as referral thresholds or authority limits. It categorizes CPVs into four types—Authority Creep, Threshold Laundering, Cumulative Sum Violation, and Context Collapse—and argues that the appropriate remedy depends on where the guarded quantity changes. To address this, the authors propose a provenance‑aware runtime architecture that evaluates policies over complete execution traces, recomputing guarded quantities from raw provenance rather than relying on step‑level outputs.

By Ashwini Kurady, Sri Sai Charith Grandhi, Rajesh Gupta, Sumit Mamoria
arXiv AI
Jul 15

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

arXiv:2607. 11951v1 Announce Type: new Abstract: Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision.

By Mohsen Arjmandi
arXiv AI
Jul 22

Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments

arXiv:2607. 18357v1 Announce Type: cross Abstract: Large language models now write a growing share of the world's code, increasingly inside agents and serving systems that compile, execute, or dispatch generated code without line-by-line review.

By Shuoming Zhang, Ruiyuan Xu, Haofeng Li, Qiuchu Yu, Yangyu Zhang, Chunwei Xia, Xiaobing Feng, Chenxi Wang, Huimin Cui, Jiacheng Zhao
arXiv AI
Sep 12

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is a benchmark that evaluates how well large language models (LLMs) understand and apply version-constraint resolution semantics, such as determining whether a version satisfies constraints like ^1.2.3 or >=2.0. The study finds that many models struggle with certain corner cases, with GPT‑5.1 performing poorly while Claude and Opus perform much better. The authors suggest that the failures stem from an activation/application gap rather than a lack of knowledge, and recommend that coding agents delegate version resolution to a dedicated resolver tool.

By Qibai Chen, Zeming Liu
arXiv Computation and Language
Aug 25

Grounded Normative Rule Generation with Structured Search

The paper introduces Grounded Normative Rule Generation (GNRS) and a new framework called GNRS-Search that uses Markov Chain Monte Carlo sampling to optimize a discrete And-Or Graph for rule synthesis. By separating operational feasibility from prose generation, the method localizes rule failures before final text creation. Evaluations on GNRS-Bench and RealCharter-Bench show significant improvements in rubric quality and executable metrics, demonstrating that the gains come from robust operational logic rather than stylistic tuning.

By Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao, Shuo Chen, Song-Chun Zhu