arXiv AI

HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule

arXiv:2606. 06679v1 Announce Type: cross Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora.

arXiv AI
Jun 18

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.

By Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang
arXiv Computation and Language
Sep 3

CLASE: A Hybrid Method for Chinese Legalese Stylistic Evaluation

CLASE is a hybrid evaluation method for Chinese legal text that combines linguistic feature-based scores with experience-guided LLM-as-a-judge scores. It learns from contrastive pairs of authentic legal documents and their LLM-generated counterparts, enabling transparent, reference-free assessment of stylistic quality. Experiments on 200 Chinese legal documents show that CLASE aligns better with human judgments than traditional metrics and offers interpretable score breakdowns and improvement suggestions.

By Yiran Rex Ma, Yuxiao Ye, Huiyuan Xie
arXiv AI
Aug 11

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.

By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv Computation and Language
Sep 1

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

The paper introduces Juris Policy Optimization (JPO), a post‑training framework designed to enhance structured legal reasoning in Chinese criminal judgment prediction. JPO first trains models with teacher‑generated rationales to guide a four‑step reasoning process, then applies reinforcement learning using a composite reward that balances prediction accuracy, reasoning completeness, and cross‑step consistency. Experiments on several open‑source language models and three Chinese legal benchmarks demonstrate that JPO consistently outperforms both supervised fine‑tuning and standard reinforcement learning baselines in terms of judgment prediction and reasoning quality.

By Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, Eric Hanchen Jiang, Jiaxin Liu, Yuan Wang, Hao Zhang, Zixia Wang, Rong Fu, Zheng Lin, Richeng Xuan, Zhichao Hu
arXiv Computation and Language
Sep 23

Mining Legal Arguments in U.S. Corporate Case Law

The paper introduces an expert‑annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations under I.R.C. §368, marking the first tree‑structured argument corpus in this domain. Each legal passage is labeled with one of five functional categories—Rule, Analysis, Conclusion, Background Facts, and Procedural History—and can be linked into directed support trees. Experiments demonstrate that functional labels are learnable and that supervised fine‑tuning improves within‑case retrieval, though cross‑case generalization remains weak.

By Luis Brena, William Jurayj, Gregory Deyesu, Zaid Al-Huneidi, Andrew Blair-Stanek, Benjamin Van Durme
arXiv Computation and Language
Aug 28

Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents

The paper introduces Legal Rule Induction (LRI), a task that seeks to extract concise, generalizable doctrinal rules from analogous judicial precedents. It presents a reproducible pipeline for constructing LRI datasets and, using Chinese law, releases the first benchmark comprising 5,121 case sets (38,088 court cases) for training and 216 expert‑annotated gold test sets. Experiments show that state‑of‑the‑art large language models struggle with over‑generalization and hallucination, but training on the new dataset significantly improves their ability to capture nuanced rule patterns across similar cases.

By Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, Yangqiu Song
arXiv Computation and Language
6d ago

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

The paper introduces an automated pipeline that converts court decisions into legal commentaries for specific German Civil Code sections, using paragraph extraction, summarization, keyword clustering, and large language models to generate headings and citation-rich sections. The system processes 4,555 decisions from the German Federal Court of Justice, evaluates the output on relevance, heading match, citation faithfulness, cluster distinction, and logical ordering, and demonstrates that rapid, low-cost commentary generation is feasible while noting limitations due to source restrictions and legal reasoning norms.

By Max Prior, Niklas Wais, Matthias Grabmair