arXiv Computation and Language

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

The paper introduces a sentence‑level benchmark for judging large language models’ ability to classify interpretive canons used by the German Federal Constitutional Court, based on Larenz’s framework. It operationalizes these canons as classification criteria, provides a dataset of court decisions annotated at the sentence level, and evaluates four LLMs with both expert hand‑written prompts and prompts optimized via Genetic‑Pareto. The results show mean F1 scores between 70.4 and 79.2, with grammatical interpretation being the easiest and systematic interpretation the hardest, and indicate that expert prompts already offer a strong baseline.

arXiv AI
Aug 11

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

arXiv:2608. 08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research.

By Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv AI
Jun 18

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.

By Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang
arXiv Computation and Language
Sep 14

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.

By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
arXiv AI
Sep 4

Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

The paper examines LLM-as-a-Judge systems used to assess AI-generated text, questioning the assumption that judgments are derived from reasoning over responses and rubrics. It finds that classifiers trained solely on rubric text can predict judge outputs, indicating that rubrics contain recoverable evaluative signals independent of the responses. Counterfactual experiments show judges often fail to adjust decisions when either the response or rubric criterion is reversed, raising doubts about the reliability of rubric-based LLM evaluation.

By Anshul Bagaria, Sowmya S Sundaram, Gokul S Krishnan, Balaraman Ravindran
arXiv Machine Learning
Sep 18

Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

The paper critiques the common practice of stopword removal in legal text analysis, showing that standard stoplists actually degrade performance on binary classification tasks involving Supreme Court opinions. By exhaustively testing the removal of each of ~18,500 candidate words, the authors find that no stoplist—generic or optimized—outperforms a no‑removal baseline, and that models cannot predict which words are beneficial to remove. The study argues that inherited preprocessing defaults can distort the doctrinal and ideological signals that legal scholars aim to recover, calling into question the validity of such practices.

By Gregory M. Dickinson