arXiv AI

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

arXiv:2608. 02621v1 Announce Type: cross Abstract: Legal benchmarks typically score final answers even when models also state legal authority.

arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv AI
Sep 24

LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law

LabourCrew is a multi‑agent Retrieval‑Augmented Generation (RAG) framework designed for trustworthy statutory question answering in labour law. It introduces three grounding mechanisms: StatuteGraph, an evidence‑exchange ledger, and a calibrated trust gate that controls false‑accept rates. Evaluated on a Bangla Labour Act QA set, LabourCrew achieves a false‑accept rate of 0.081 and higher answer relevancy than existing RAG methods, demonstrating that calibrated abstention is key to auditable legal QA.

By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
arXiv Computation and Language
Sep 17

Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.

By Maksym Taranukhin, Vered Shwartz
arXiv Computation and Language
Sep 24

Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction

The paper presents a cross‑lingual legal QA system for Vietnamese labour law, introducing a bilingual evaluation suite of 231 Vietnamese–English question–answer pairs, 75 of which are annotated for five complex legal reasoning phenomena. It evaluates a verifier‑guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions, and introduces six automatic diagnostics for faithfulness to retrieved evidence. Experiments show that dense retrieval outperforms sparse and hybrid retrieval, translation placement has no significant effect on diagnostics, and verifier‑guided correction modestly improves citation preservation but not other dimensions, with human evaluation indicating a gap between automatic diagnostics and human judgments.

By Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson
arXiv AI
Jun 18

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.

By Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang