arXiv:2607. 04261v1 Announce Type: new Abstract: Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting.
By Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin, M{\aa}ns Magnusson, Huiyuan Xie, Hao Tian Yeung, Felix Steffek
arXiv:2606. 06679v1 Announce Type: cross Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora.
By Xi Xuan, Wenxin Zhang, Yufei Zhou, King-kui Sin, Chunyu Kit
Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.
By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
The paper evaluates legal text classification models for Korean sexual offense cases, comparing traditional machine learning, large language models, and fine‑tuned domain models. Fine‑tuned KLUE‑BERT achieved the highest accuracy of 99.3%, outperforming GPT‑3.5, GPT‑4.0, and other traditional approaches. Explainable AI techniques were used to analyze predictions, revealing linguistic features that influence decisions and highlighting limitations in capturing subtle textual cues, especially in real‑world KICS data.
By Jeongmin Lee
arXiv:2609.22529v1 Announce Type: new
Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its t...
By Genis Skura, Roland Bouffanais, Didier Wernli
arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.
By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv:2606. 23716v1 Announce Type: cross Abstract: Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights.
By Andrew Lou, David Shin
The paper introduces a sentence‑level benchmark for judging large language models’ ability to classify interpretive canons used by the German Federal Constitutional Court, based on Larenz’s framework. It operationalizes these canons as classification criteria, provides a dataset of court decisions annotated at the sentence level, and evaluates four LLMs with both expert hand‑written prompts and prompts optimized via Genetic‑Pareto. The results show mean F1 scores between 70.4 and 79.2, with grammatical interpretation being the easiest and systematic interpretation the hardest, and indicate that expert prompts already offer a strong baseline.
By Felix Ringe
arXiv:2608.20391v1 Announce Type: new
Abstract: Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of...
By Amirhossein Afsharrad, Seyed Shahabeddin Mousavi
arXiv:2604. 04790v2 Announce Type: replace-cross Abstract: Natural language processing (NLP) advances have powered a generation of LegalTech systems, but Turkish law remains under-served by domain-specific data and models.
By Mehmet Utku \"Ozt\"urk, Tansu T\"urko\u{g}lu, Buse Buz-Yalug
arXiv:2607. 03325v1 Announce Type: cross Abstract: We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism.
By Giovanni Piccioli, Alessia Fidelangeli, Piera Santin, Pierpaolo Vivo
The paper introduces Legal Rule Induction (LRI), a task that seeks to extract concise, generalizable doctrinal rules from analogous judicial precedents. It presents a reproducible pipeline for constructing LRI datasets and, using Chinese law, releases the first benchmark comprising 5,121 case sets (38,088 court cases) for training and 216 expert‑annotated gold test sets. Experiments show that state‑of‑the‑art large language models struggle with over‑generalization and hallucination, but training on the new dataset significantly improves their ability to capture nuanced rule patterns across similar cases.
By Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, Yangqiu Song