arXiv AI

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

arXiv:2606. 07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational risks.

arXiv Computation and Language
Aug 25

Grounded Normative Rule Generation with Structured Search

The paper introduces Grounded Normative Rule Generation (GNRS) and a new framework called GNRS-Search that uses Markov Chain Monte Carlo sampling to optimize a discrete And-Or Graph for rule synthesis. By separating operational feasibility from prose generation, the method localizes rule failures before final text creation. Evaluations on GNRS-Bench and RealCharter-Bench show significant improvements in rubric quality and executable metrics, demonstrating that the gains come from robust operational logic rather than stylistic tuning.

By Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao, Shuo Chen, Song-Chun Zhu
arXiv AI
Sep 17

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.

By Mika Okamoto, Ansel Kaplan Erol
arXiv AI
Jul 23

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

arXiv:2607. 19865v1 Announce Type: new Abstract: As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows.

By Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun