The paper evaluates legal text classification models for Korean sexual offense cases, comparing traditional machine learning, large language models, and fine‑tuned domain models. Fine‑tuned KLUE‑BERT achieved the highest accuracy of 99.3%, outperforming GPT‑3.5, GPT‑4.0, and other traditional approaches. Explainable AI techniques were used to analyze predictions, revealing linguistic features that influence decisions and highlighting limitations in capturing subtle textual cues, especially in real‑world KICS data.
By Jeongmin Lee
arXiv:2605. 29738v2 Announce Type: replace-cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.
By Volodymyr Ovcharov
arXiv:2609.16010v1 Announce Type: new
Abstract: The complexity of legal language and limited accessibility to legal information pose significant challenges to justice delivery in Nepal. Traditional l...
By Ranjit Raut, Tishya Dhakal, Aaryan Shakya, Bhabuk Thapa, Prasiddha Koirala, Bal Krishna Bal
arXiv:2609.22529v1 Announce Type: new
Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its t...
By Genis Skura, Roland Bouffanais, Didier Wernli
arXiv:2607. 29066v1 Announce Type: cross Abstract: Deception detection has critical implications for legal proceedings, law enforcement, and online security.
By Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb
The paper surveys how large language models (LLMs) are being applied in legal tasks such as judgement prediction, document analysis, and drafting. It reviews the benefits of automation while highlighting legal challenges like privacy, bias, and explainability. The authors also discuss data resources for legal domain specialization and outline future research directions.
By Zhongxiang Sun
arXiv:2606. 23716v1 Announce Type: cross Abstract: Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who cannot access lawyers in order to understand and exercise their legal rights.
By Andrew Lou, David Shin
arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.
By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv:2012. 02110v2 Announce Type: replace-cross Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa.
By Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, Martin Boeker
The paper critiques the common practice of stopword removal in legal text analysis, showing that standard stoplists actually degrade performance on binary classification tasks involving Supreme Court opinions. By exhaustively testing the removal of each of ~18,500 candidate words, the authors find that no stoplist—generic or optimized—outperforms a no‑removal baseline, and that models cannot predict which words are beneficial to remove. The study argues that inherited preprocessing defaults can distort the doctrinal and ideological signals that legal scholars aim to recover, calling into question the validity of such practices.
By Gregory M. Dickinson
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
The paper introduces Legal Rule Induction (LRI), a task that seeks to extract concise, generalizable doctrinal rules from analogous judicial precedents. It presents a reproducible pipeline for constructing LRI datasets and, using Chinese law, releases the first benchmark comprising 5,121 case sets (38,088 court cases) for training and 216 expert‑annotated gold test sets. Experiments show that state‑of‑the‑art large language models struggle with over‑generalization and hallucination, but training on the new dataset significantly improves their ability to capture nuanced rule patterns across similar cases.
By Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, Yangqiu Song