arXiv AI

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

arXiv:2606. 18021v1 Announce Type: new Abstract: AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment.

arXiv AI
Sep 11

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

LexAgentHallu is a new benchmark that profiles hallucinations in legal agents across multi-step interactions. It contains 3,414 instances spanning 17 legal categories and 6 task types, each annotated with a dual-layer taxonomy of 7 high-level and 27 fine-grained hallucination categories. The benchmark introduces fine-grained metrics to quantify and localize failures along an agent’s execution path, revealing patterns such as the Right-Answer-Wrong-Reason effect and clustered hallucination subclasses.

By Yujin Zhou, Mingxuan Zheng, Chuxue Cao, Huang Yidan, Jiale Chen, Yike Guo, Sirui Han
arXiv AI
Jun 4

Parthenon Law: A Self-Evolving Legal-Agent Framework

arXiv:2606. 04602v1 Announce Type: new Abstract: As agents grow more capable, legal-domain LLM agents promise to turn document-heavy matters into reviewable work products -- yet reliable deployment faces three obstacles: no large-scale evidence on how today's strongest model-and-harness combinations behave on end-to-end legal matters; no agent architecture adapted to the legal vertical, only general-purpose harnesses; and, in a setting that keeps shifting with new facts, authorities, and deadlines, no mechanism for systems to learn from their own outcomes.

By Hejia Geng, Leo Liu
arXiv AI
Sep 18

Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems

The paper introduces Governance-as-Code (GaC), a framework that translates the EU AI Act’s technical requirements into 43 machine‑checkable acceptance criteria across six compliance modules. GaC runs within a CI/CD pipeline, producing Article‑indexed audit evidence and providing actual Rego policy code. The authors validate GaC on two enterprise deployments, showing it reproduces manual audit findings—including three penalty‑triggering violations—while reducing audit labor by about 75%.

By Rudrendu Kumar Paul, Sourav Nandy
Hugging Face Trending Papers
Jun 3

Parthenon Law: A Self-Evolving Legal-Agent Framework

As agents grow more capable, legal-domain LLM agents promise to turn document-heavy matters into reviewable work products -- yet reliable deployment faces three obstacles: no large-scale evidence on how today's strongest model-and-harness combinations behave on end-to-end legal matters; no agent architecture adapted to the legal vertical, only general-purpose harnesses; and, in a setting that keeps shifting with new facts, authorities, and deadlines, no mechanism for systems to learn from their own outcomes. We address each.

arXiv Computation and Language
Sep 17

Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.

By Maksym Taranukhin, Vered Shwartz
arXiv AI
3d ago

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.

By Serhii Zabolotnii
arXiv AI
Aug 19

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.

By Adam Mazzocchetti