arXiv:2608. 12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation.
By Mika Okamoto, Ansel Kaplan Erol, Kutluhan Erol
arXiv:2606. 02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceeded at all.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2606. 02965v2 Announce Type: replace Abstract: As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2606. 04455v1 Announce Type: new Abstract: Current AI benchmarks evaluate agents on task execution within human-designed workflows.
By Xinyu Lu, Tianshu Wang, Pengbo Wang, zujie wen, Zhiqiang Zhang, Jun Zhou, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
arXiv:2609.39107v1 Announce Type: new
Abstract: Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and...
By Yan Zhang, Chuming Wei, Ruien Li, Yaoyao Peng, Wusheng Zhang, Guangwen Yang
The paper introduces Grounded Normative Rule Generation (GNRS) and a new framework called GNRS-Search that uses Markov Chain Monte Carlo sampling to optimize a discrete And-Or Graph for rule synthesis. By separating operational feasibility from prose generation, the method localizes rule failures before final text creation. Evaluations on GNRS-Bench and RealCharter-Bench show significant improvements in rubric quality and executable metrics, demonstrating that the gains come from robust operational logic rather than stylistic tuning.
By Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao, Shuo Chen, Song-Chun Zhu