Hugging Face Trending Papers

Mining Legal Arguments in U.S. Corporate Case Law

arXiv Computation and Language
Sep 23

Mining Legal Arguments in U.S. Corporate Case Law

The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.

By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv Computation and Language
4d ago

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

The paper introduces an automated pipeline that converts court decisions into legal commentaries for specific German Civil Code sections, using paragraph extraction, summarization, keyword clustering, and large language models to generate headings and citation-rich sections. The system processes 4,555 decisions from the German Federal Court of Justice, evaluates the output on relevance, heading match, citation faithfulness, cluster distinction, and logical ordering, and demonstrates that rapid, low-cost commentary generation is feasible while noting limitations due to source restrictions and legal reasoning norms.

By Max Prior, Niklas Wais, Matthias Grabmair
arXiv AI
Aug 11

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.

By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv Computation and Language
Sep 4

LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

LexIssue introduces a benchmark for identifying disputed legal issues in Chinese civil litigation, comprising 430 real‑world cases and 1,303 expert‑annotated issues. The dataset is built around a hierarchical schema that links free‑form issue descriptions to structured legal categories, enabling two complementary tasks: issue generation and issue classification. A retrieval‑augmented knowledge base covering 27 causes of action and 441 issue entries is provided, and experiments show that incorporating this knowledge consistently improves model performance on the tasks.

By Huiyuan Xie, Yuqin Huang, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye
arXiv Computation and Language
Sep 24

LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning

LEGO is a dual‑module framework that combines a Legal Expert GraphRAG system with an expert Chain‑of‑Thought approach to enhance complex legal reasoning. The GraphRAG component uses an expert‑annotated civil code graph and a greedy normative‑coverage retrieval algorithm to extract relevant provision subgraphs, while the Chain‑of‑Thought module structures retrieved provisions and case facts into a Provision‑Fact‑Conclusion reasoning flow. Using a Qwen3‑8B backbone, LEGO achieves 40.53% exact‑match accuracy on LawExamQA_Civil, surpassing baseline RAG and CoT models and matching larger models on multi‑hop and open‑ended benchmarks, with ablation studies confirming the complementary benefits of both modules.

By Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng, Yukun Yan, Zhi Zheng, Antonino Rotolo, Yun Liu, Weixing Shen
arXiv AI
Sep 18

By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode

The paper introduces a systematic method for comparing different formalizations of the same legal provision by analyzing their inferences on individual cases. It matches formalizations at the node level, derives shared interfaces, and uses a SAT solver to identify edge cases where any two formalizations disagree. The authors apply this approach to ten EU provisions formalized by nine advanced LLMs, finding that behavioral divergence is largely uncorrelated with structural agreement and that the resulting edge cases expose distinct types of disagreement, some reflecting real legal controversies.

By Julius Vernie, Matthias Grabmair
Hugging Face Trending Papers
Jul 13

Generative Chinese Statute Retrieval

Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries and formal statutory language. In this paper, we propose GCSR, a generative statute retrieval framework that reformulates statute retrieval as a sequence generation problem and internalizes statutory knowledge into a generative model.