arXiv AI

Guidelines for the Annotation and Visualization of Legal Argumentation Structures in Chinese Judicial Decisions

arXiv:2603. 05171v2 Announce Type: replace-cross Abstract: This Guideline presents a systematic and operationalizable annotation framework for representing legal argumentation structures in judicial decisions.

arXiv Computation and Language
Sep 24

LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning

LEGO is a dual‑module framework that combines a Legal Expert GraphRAG system with an expert Chain‑of‑Thought approach to enhance complex legal reasoning. The GraphRAG component uses an expert‑annotated civil code graph and a greedy normative‑coverage retrieval algorithm to extract relevant provision subgraphs, while the Chain‑of‑Thought module structures retrieved provisions and case facts into a Provision‑Fact‑Conclusion reasoning flow. Using a Qwen3‑8B backbone, LEGO achieves 40.53% exact‑match accuracy on LawExamQA_Civil, surpassing baseline RAG and CoT models and matching larger models on multi‑hop and open‑ended benchmarks, with ablation studies confirming the complementary benefits of both modules.

By Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng, Yukun Yan, Zhi Zheng, Antonino Rotolo, Yun Liu, Weixing Shen
arXiv Computation and Language
Sep 23

Mining Legal Arguments in U.S. Corporate Case Law

The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.

By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv AI
Sep 18

By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode

The paper introduces a systematic method for comparing different formalizations of the same legal provision by analyzing their inferences on individual cases. It matches formalizations at the node level, derives shared interfaces, and uses a SAT solver to identify edge cases where any two formalizations disagree. The authors apply this approach to ten EU provisions formalized by nine advanced LLMs, finding that behavioral divergence is largely uncorrelated with structural agreement and that the resulting edge cases expose distinct types of disagreement, some reflecting real legal controversies.

By Julius Vernie, Matthias Grabmair
arXiv Computation and Language
Sep 1

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

The paper introduces Juris Policy Optimization (JPO), a post‑training framework designed to enhance structured legal reasoning in Chinese criminal judgment prediction. JPO first trains models with teacher‑generated rationales to guide a four‑step reasoning process, then applies reinforcement learning using a composite reward that balances prediction accuracy, reasoning completeness, and cross‑step consistency. Experiments on several open‑source language models and three Chinese legal benchmarks demonstrate that JPO consistently outperforms both supervised fine‑tuning and standard reinforcement learning baselines in terms of judgment prediction and reasoning quality.

By Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, Eric Hanchen Jiang, Jiaxin Liu, Yuan Wang, Hao Zhang, Zixia Wang, Rong Fu, Zheng Lin, Richeng Xuan, Zhichao Hu
arXiv Computation and Language
6d ago

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

The paper introduces an automated pipeline that converts court decisions into legal commentaries for specific German Civil Code sections, using paragraph extraction, summarization, keyword clustering, and large language models to generate headings and citation-rich sections. The system processes 4,555 decisions from the German Federal Court of Justice, evaluates the output on relevance, heading match, citation faithfulness, cluster distinction, and logical ordering, and demonstrates that rapid, low-cost commentary generation is feasible while noting limitations due to source restrictions and legal reasoning norms.

By Max Prior, Niklas Wais, Matthias Grabmair
arXiv AI
Aug 19

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.

By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen