arXiv Computation and Language

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

The paper introduces an automated pipeline that converts court decisions into legal commentaries for specific German Civil Code sections, using paragraph extraction, summarization, keyword clustering, and large language models to generate headings and citation-rich sections. The system processes 4,555 decisions from the German Federal Court of Justice, evaluates the output on relevance, heading match, citation faithfulness, cluster distinction, and logical ordering, and demonstrates that rapid, low-cost commentary generation is feasible while noting limitations due to source restrictions and legal reasoning norms.

arXiv Computation and Language
Sep 23

Mining Legal Arguments in U.S. Corporate Case Law

The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.

By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv AI
Aug 11

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.

By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv AI
Sep 18

By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode

The paper introduces a systematic method for comparing different formalizations of the same legal provision by analyzing their inferences on individual cases. It matches formalizations at the node level, derives shared interfaces, and uses a SAT solver to identify edge cases where any two formalizations disagree. The authors apply this approach to ten EU provisions formalized by nine advanced LLMs, finding that behavioral divergence is largely uncorrelated with structural agreement and that the resulting edge cases expose distinct types of disagreement, some reflecting real legal controversies.

By Julius Vernie, Matthias Grabmair
arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu