LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
LexReward is a taxonomy-driven reward framework designed for legal language models, evaluating responses across three dimensions: Style (lexical and syntactic quality), Element (legal subjects, facts, statutes, and decisions), and Chain (order, completeness, correctness, and non-redundancy of reasoning). The framework introduces rubrics that define evaluation criteria and quality levels for each dimension, generating pairwise preference data used for Direct Preference Optimization (DPO) and reward-model training. Experiments demonstrate that rubric-based rewards effectively differentiate legal responses of varying quality, and that DPO training improves performance across all dimensions; the resulting reward models (LexRM) enable reinforcement learning to enhance policy performance in each specific dimension without needing reference answers.
arXiv:2609.14739v1 Announce Type: cross Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and a...
The paper introduces Juris Policy Optimization (JPO), a post‑training framework designed to enhance structured legal reasoning in Chinese criminal judgment prediction. JPO first trains models with teacher‑generated rationales to guide a four‑step reasoning process, then applies reinforcement learning using a composite reward that balances prediction accuracy, reasoning completeness, and cross‑step consistency. Experiments on several open‑source language models and three Chinese legal benchmarks demonstrate that JPO consistently outperforms both supervised fine‑tuning and standard reinforcement learning baselines in terms of judgment prediction and reasoning quality.
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.
arXiv:2607. 19181v1 Announce Type: cross Abstract: Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires.