The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.
By Maksym Taranukhin, Vered Shwartz
arXiv:2606. 23913v1 Announce Type: new Abstract: This article develops an architecture that creates a formally verifiable reward signal to train legal AI, adapting the LLM proposes, verifier disposes paradigm from mathematical AI to the distinctive demands of law.
By Armin Heydari (Harvard University), Torben Leowald (Columbia University)
arXiv:2606. 16118v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on reasoning tasks, but whether this reflects faithful logical inference or heuristic approximation remains unclear.
By Olivia Peiyu Wang, Sanna Wong-Toropainen, Daneshvar Amrollahi, Ryan Bai, Tashvi Bansal, Arush Garg, Leilani H. Gilpin
LEGO is a dual‑module framework that combines a Legal Expert GraphRAG system with an expert Chain‑of‑Thought approach to enhance complex legal reasoning. The GraphRAG component uses an expert‑annotated civil code graph and a greedy normative‑coverage retrieval algorithm to extract relevant provision subgraphs, while the Chain‑of‑Thought module structures retrieved provisions and case facts into a Provision‑Fact‑Conclusion reasoning flow. Using a Qwen3‑8B backbone, LEGO achieves 40.53% exact‑match accuracy on LawExamQA_Civil, surpassing baseline RAG and CoT models and matching larger models on multi‑hop and open‑ended benchmarks, with ablation studies confirming the complementary benefits of both modules.
By Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng, Yukun Yan, Zhi Zheng, Antonino Rotolo, Yun Liu, Weixing Shen
The paper introduces Legal Rule Induction (LRI), a task that seeks to extract concise, generalizable doctrinal rules from analogous judicial precedents. It presents a reproducible pipeline for constructing LRI datasets and, using Chinese law, releases the first benchmark comprising 5,121 case sets (38,088 court cases) for training and 216 expert‑annotated gold test sets. Experiments show that state‑of‑the‑art large language models struggle with over‑generalization and hallucination, but training on the new dataset significantly improves their ability to capture nuanced rule patterns across similar cases.
By Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, Yangqiu Song
arXiv:2608. 14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination.
By Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.
By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
arXiv:2603. 05171v2 Announce Type: replace-cross Abstract: This Guideline presents a systematic and operationalizable annotation framework for representing legal argumentation structures in judicial decisions.
By Kun Chen, Xianglei Liao, Kaixue Fei, Yi Xing, Xinrui Li
arXiv:2607. 03325v1 Announce Type: cross Abstract: We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism.
By Giovanni Piccioli, Alessia Fidelangeli, Piera Santin, Pierpaolo Vivo
arXiv:2608. 02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation.
By Benjamin Fresz, Elena Dubovitskaya, Marco F. Huber
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
arXiv:2606. 18158v1 Announce Type: cross Abstract: Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure.
By Mich\`ele Finck