LexReward is a taxonomy-driven reward framework designed for legal language models, evaluating responses across three dimensions: Style (lexical and syntactic quality), Element (legal subjects, facts, statutes, and decisions), and Chain (order, completeness, correctness, and non-redundancy of reasoning). The framework introduces rubrics that define evaluation criteria and quality levels for each dimension, generating pairwise preference data used for Direct Preference Optimization (DPO) and reward-model training. Experiments demonstrate that rubric-based rewards effectively differentiate legal responses of varying quality, and that DPO training improves performance across all dimensions; the resulting reward models (LexRM) enable reinforcement learning to enhance policy performance in each specific dimension without needing reference answers.
By Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu
arXiv:2609.14739v1 Announce Type: cross
Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and a...
By Rilton Franzone, Valentin No\"el, Puyu Wang, Philip Torr, Fabio J. Fehr
The paper introduces Juris Policy Optimization (JPO), a post‑training framework designed to enhance structured legal reasoning in Chinese criminal judgment prediction. JPO first trains models with teacher‑generated rationales to guide a four‑step reasoning process, then applies reinforcement learning using a composite reward that balances prediction accuracy, reasoning completeness, and cross‑step consistency. Experiments on several open‑source language models and three Chinese legal benchmarks demonstrate that JPO consistently outperforms both supervised fine‑tuning and standard reinforcement learning baselines in terms of judgment prediction and reasoning quality.
By Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, Eric Hanchen Jiang, Jiaxin Liu, Yuan Wang, Hao Zhang, Zixia Wang, Rong Fu, Zheng Lin, Richeng Xuan, Zhichao Hu
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.
By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv:2607. 19181v1 Announce Type: cross Abstract: Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires.
By Aixiu An, Michael Jungo, Eloi Eynard, Mark Drenhaus, Andreas Fischer, Jean Hennebert, S\'ebastien Rumley
The paper surveys how large language models (LLMs) are being applied in legal tasks such as judgement prediction, document analysis, and drafting. It reviews the benefits of automation while highlighting legal challenges like privacy, bias, and explainability. The authors also discuss data resources for legal domain specialization and outline future research directions.
By Zhongxiang Sun
arXiv:2608. 14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination.
By Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
arXiv:2605.25920v2 Announce Type: replace
Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constra...
By Wei Fan, Yining Zhou, Mufan Zhang, Yanbing Weng, Yiran HU, Tianshi Zheng, Baixuan Xu, Chunyang Li, Jianhui Yang, Haoran Li, Yangqiu Song
arXiv:2605. 25240v2 Announce Type: replace-cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs.
By Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, Julian Nyarko
Legal argument mining supports passage classification, retrieval, and argument completion. This work introduces an expert-annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations...
arXiv:2606. 23913v1 Announce Type: new Abstract: This article develops an architecture that creates a formally verifiable reward signal to train legal AI, adapting the LLM proposes, verifier disposes paradigm from mathematical AI to the distinctive demands of law.
By Armin Heydari (Harvard University), Torben Leowald (Columbia University)