arXiv AI By Yiyang Zhao, Zhuo Zhang, Qingxuan Le, Lizhen Qu, Zenglin Xu

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

Read the original on arXiv AI →

arXiv:2606. 07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational risks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Grounded Normative Rule Generation with Structured Search

The paper introduces Grounded Normative Rule Generation (GNRS) and a new framework called GNRS-Search that uses Markov Chain Monte Carlo sampling to optimize a discrete And-Or Graph for rule synthesis. By separating operational feasibility from prose generation, the method localizes rule failures before final text creation. Evaluations on GNRS-Bench and RealCharter-Bench show significant improvements in rubric quality and executable metrics, demonstrating that the gains come from robust operational logic rather than stylistic tuning.

By Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao, Shuo Chen, Song-Chun Zhu