arXiv:2606.09389v2 Announce Type: replace
Abstract: As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses...
By Yifan Chen, Haitao Li, Yiran Hu, Kaisong Song, Jun Lin, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu
arXiv:2606. 06679v1 Announce Type: cross Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora.
By Xi Xuan, Wenxin Zhang, Yufei Zhou, King-kui Sin, Chunyu Kit
Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.
By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv:2605. 28183v2 Announce Type: replace-cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law.
By Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Ko\c{c}ak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
CLASE is a hybrid evaluation method for Chinese legal text that combines linguistic feature-based scores with experience-guided LLM-as-a-judge scores. It learns from contrastive pairs of authentic legal documents and their LLM-generated counterparts, enabling transparent, reference-free assessment of stylistic quality. Experiments on 200 Chinese legal documents show that CLASE aligns better with human judgments than traditional metrics and offers interpretable score breakdowns and improvement suggestions.
By Yiran Rex Ma, Yuxiao Ye, Huiyuan Xie
CoAL‑RAG is a complexity‑aware legal retrieval‑augmented generation method that adapts its retrieval strategy based on a multi‑dimensional evaluation of question essence and retrieval consistency. It quantifies reasoning demand from the logical structure of a question and uses the discrepancy between semantic and keyword retrieval to gauge problem complexity, thereby selecting the most suitable retrieval approach and filtering context dynamically. Experiments show that CoAL‑RAG outperforms baseline models on Chinese legal benchmarks (SocialLawQA, LawBench) with a 42.5% BLEU improvement and 3.6× ROUGE‑L, while also achieving strong cross‑jurisdictional performance on English datasets (LexGLUE, CaseHold).
By Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
arXiv:2605. 29738v2 Announce Type: replace-cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.
By Volodymyr Ovcharov
The paper introduces Juris Policy Optimization (JPO), a post‑training framework designed to enhance structured legal reasoning in Chinese criminal judgment prediction. JPO first trains models with teacher‑generated rationales to guide a four‑step reasoning process, then applies reinforcement learning using a composite reward that balances prediction accuracy, reasoning completeness, and cross‑step consistency. Experiments on several open‑source language models and three Chinese legal benchmarks demonstrate that JPO consistently outperforms both supervised fine‑tuning and standard reinforcement learning baselines in terms of judgment prediction and reasoning quality.
By Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei, Chenghua Zhu, Junhao Gong, Jiachen Qian, Eric Hanchen Jiang, Jiaxin Liu, Yuan Wang, Hao Zhang, Zixia Wang, Rong Fu, Zheng Lin, Richeng Xuan, Zhichao Hu
The paper introduces Legal Rule Induction (LRI), a task that seeks to extract concise, generalizable doctrinal rules from analogous judicial precedents. It presents a reproducible pipeline for constructing LRI datasets and, using Chinese law, releases the first benchmark comprising 5,121 case sets (38,088 court cases) for training and 216 expert‑annotated gold test sets. Experiments show that state‑of‑the‑art large language models struggle with over‑generalization and hallucination, but training on the new dataset significantly improves their ability to capture nuanced rule patterns across similar cases.
By Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, Yangqiu Song
The paper presents a smartphone‑compatible, retrieval‑augmented language model tailored to Bangladeshi statutory law. By distilling a 9‑billion‑parameter Gemma‑2 teacher into a 2‑billion‑parameter student using supervised fine‑tuning and QLoRA, the authors achieve significant gains in ROUGE‑L and BERTScore on an English benchmark while keeping the model lightweight (1.6 GB) and operable offline on a Pixel 6. The system retrieves from 36,029 statutory passages using a hybrid dense/BM25 approach, and cross‑lingual evaluation shows effective Bangla query handling against an English‑only corpus, with a practicing lawyer rating the responses highly in a pilot study.
By MD. Nafis Kamal, Mahadi Hasan Fahim, Talha Ridwan, Nadifa Zaman, Fariha Roushon Florin, Farig Yousuf Sadeque, Saadat Rafid Ahmed
arXiv:2607. 18825v1 Announce Type: cross Abstract: This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the Indian legal context.
By Shubham Kumar Nigam, Shubham Kumar Mishra, Noel Shallum, Kripabandhu Ghosh, Arnab Bhattacharya