arXiv AI

Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines

arXiv:2607. 05680v1 Announce Type: cross Abstract: AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especially when that advice is legally correct but socially controversial.

arXiv AI
Sep 10

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv AI
Aug 19

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.

By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
arXiv AI
Aug 6

Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools

arXiv:2604. 26233v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions.

By Oisin Suttle, David Lillis
arXiv Computation and Language
Sep 17

Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.

By Maksym Taranukhin, Vered Shwartz
arXiv AI
Aug 24

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?

The study examines how Large Language Models (LLMs) can exhibit an ‘inertia of confidence’, giving incorrect legal verdicts with high certainty, and tests this on Indian Contract Act cases. Phase I audits ChatGPT, Meta AI, and Perplexity AI, introducing the High‑Confidence Error Rate (HCER) to measure dangerous certainty, finding Meta AI most prone to errors. Phase II surveys 380 Indian law students, revealing that exposure to hallucinated citations increases verification efforts but most students lack formal ethical AI training.

By Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
arXiv AI
Sep 7

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.

By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert