The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.
By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.
By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv:2608.21409v1 Announce Type: cross
Abstract: In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent...
By Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.
By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv:2609.23083v1 Announce Type: new
Abstract: The distinction between the spirit and letter of the law is a central issue across research and everyday life, and a growing concern for building safe,...
By Peng Qian, Andrew Li, Sam Chen, Sonia K. Murthy, Yonatan Belinkov, Tomer D. Ullman