arXiv AI By Mich\`ele Finck

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

Read the original on arXiv AI →

arXiv:2606. 18158v1 Announce Type: cross Abstract: Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode

The paper introduces a systematic method for comparing different formalizations of the same legal provision by analyzing their inferences on individual cases. It matches formalizations at the node level, derives shared interfaces, and uses a SAT solver to identify edge cases where any two formalizations disagree. The authors apply this approach to ten EU provisions formalized by nine advanced LLMs, finding that behavioral divergence is largely uncorrelated with structural agreement and that the resulting edge cases expose distinct types of disagreement, some reflecting real legal controversies.

By Julius Vernie, Matthias Grabmair
arXiv AI
Aug 19

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.

By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen