The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.
By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv:2606. 00013v1 Announce Type: cross Abstract: Social conformity is a well-documented phenomenon in which individuals shift their opinions towards those of a social majority.
By Yana Venerina, Dmitry Koch, Nare Meloyan, Gerda Prutko, Valeriia Lelik, Victoria Taova, Andrey Kurpatov
The study examines whether large language models (LLMs) can perform legally meaningful reasoning by testing OpenAI GPT 5.4 on European Court of Human Rights case forecasting. Using various prompting strategies, the authors find that the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet only weakly aligned with human annotators. The expert-curated prompt yields more comprehensive reasoning but does not improve prediction accuracy, leading the authors to caution against relying solely on automated LLM evaluation or using task accuracy as a proxy for reasoning quality.
By Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
arXiv:2603.19042v5 Announce Type: replace
Abstract: The integration of artificial intelligence (AI) into judicial decision making -- particularly in pretrial, sentencing, and parole contexts -- has g...
By Arthur Dyevre, Ahmad Shahvaroughi
arXiv:2508. 07872v2 Announce Type: replace-cross Abstract: Uncertainty in artificial intelligence (AI) predictions raises pressing legal and ethical questions for AI-assisted decision-making.
By Holli Sargeant, Mackenzie Jorgensen, Arina Shah, Sam Goring, Adrian Weller, Umang Bhatt
arXiv:2604. 26233v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions.
By Oisin Suttle, David Lillis
arXiv:2608. 05602v1 Announce Type: new Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled.
By Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt
The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.
By Maksym Taranukhin, Vered Shwartz
arXiv:2608.29453v1 Announce Type: cross
Abstract: As AI becomes increasingly integrated into clinical practice, it is playing a growing role in medical decision making. Medicine, however, is a high s...
By Jiayuan Zhu, Jiazhen Pan, Fenglin Liu, Minhao Hu, Junde Wu
arXiv:2608.28593v1 Announce Type: new
Abstract: With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in...
By Cindy Delage, St\'ephane Canu, Marc D\'ecombas, Jonathan Foureur
The study examines how Large Language Models (LLMs) can exhibit an ‘inertia of confidence’, giving incorrect legal verdicts with high certainty, and tests this on Indian Contract Act cases. Phase I audits ChatGPT, Meta AI, and Perplexity AI, introducing the High‑Confidence Error Rate (HCER) to measure dangerous certainty, finding Meta AI most prone to errors. Phase II surveys 380 Indian law students, revealing that exposure to hallucinated citations increases verification efforts but most students lack formal ethical AI training.
By Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.
By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert