Legal argument mining supports passage classification, retrieval, and argument completion. This work introduces an expert-annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations...
The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.
By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv:2608.03756v2 Announce Type: cross
Abstract: A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pi...
By Theresia Veronika Rampisela, Henrik Palmer Olsen, Giovanni Colavizza
arXiv:2607. 03325v1 Announce Type: cross Abstract: We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism.
By Giovanni Piccioli, Alessia Fidelangeli, Piera Santin, Pierpaolo Vivo
arXiv:2603. 22973v2 Announce Type: replace Abstract: Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity.
By Avrile Floro (UPHF), Tamara Dhorasoo (UPHF), Soline Pellez (UPHF), Nils Holzenberger
arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.
By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv:2606. 06679v1 Announce Type: cross Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora.
By Xi Xuan, Wenxin Zhang, Yufei Zhou, King-kui Sin, Chunyu Kit
arXiv:2607. 09094v1 Announce Type: cross Abstract: Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research.
By Devanshu Verma, Vasudha Bhatnagar, Vikas Kumar, Balaji Ganesan
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
The paper introduces a systematic method for comparing different formalizations of the same legal provision by analyzing their inferences on individual cases. It matches formalizations at the node level, derives shared interfaces, and uses a SAT solver to identify edge cases where any two formalizations disagree. The authors apply this approach to ten EU provisions formalized by nine advanced LLMs, finding that behavioral divergence is largely uncorrelated with structural agreement and that the resulting edge cases expose distinct types of disagreement, some reflecting real legal controversies.
By Julius Vernie, Matthias Grabmair
arXiv:2505. 02763v2 Announce Type: replace-cross Abstract: One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without calling for much discretion.
By Matthew Dahl, Eric Mart\'inez
The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.
By Yao Dou, Benjamin Mamut, Wei Xu