arXiv AI By Linh Le, Melanie Bui, My Chiffon Nguyen, Zachary Schlosser, David Williams-King

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

Read the original on arXiv AI →

arXiv:2609. 03553v1 Announce Type: new Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

Parthenon Law: A Self-Evolving Legal-Agent Framework

arXiv:2606. 04602v1 Announce Type: new Abstract: As agents grow more capable, legal-domain LLM agents promise to turn document-heavy matters into reviewable work products -- yet reliable deployment faces three obstacles: no large-scale evidence on how today's strongest model-and-harness combinations behave on end-to-end legal matters; no agent architecture adapted to the legal vertical, only general-purpose harnesses; and, in a setting that keeps shifting with new facts, authorities, and deadlines, no mechanism for systems to learn from their own outcomes.

By Hejia Geng, Leo Liu