arXiv AI

Man and machine: artificial intelligence and judicial decision making

arXiv Machine Learning
Aug 12

Do Judges Behave Like Algorithms?

arXiv:2608. 10400v1 Announce Type: new Abstract: What if judges already behave like algorithms?

By Riya Manchanda, Eric Chen, Chloe Zhu, Cynthia Rudin, Brandon Garrett, Songman Kang
arXiv AI
Sep 10

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv AI
Aug 24

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?

The study examines how Large Language Models (LLMs) can exhibit an ‘inertia of confidence’, giving incorrect legal verdicts with high certainty, and tests this on Indian Contract Act cases. Phase I audits ChatGPT, Meta AI, and Perplexity AI, introducing the High‑Confidence Error Rate (HCER) to measure dangerous certainty, finding Meta AI most prone to errors. Phase II surveys 380 Indian law students, revealing that exposure to hallucinated citations increases verification efforts but most students lack formal ethical AI training.

By Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
arXiv AI
Sep 3

OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction

The paper introduces OBJECTION, an inference-time pipeline that adds an Adversarial Lawyer Agent to each of the three reasoning steps—offense, unlawfulness, and culpability—in legal judgment prediction models. By actively injecting defense arguments, the agent challenges the model’s default assumption of guilt, which is common in datasets biased toward guilty outcomes. Using a new Natural Innocent dataset of 3.4k real cases, OBJECTION reduces the False Guilty Rate from 82.93% to 16.69%, demonstrating significant improvement in substantive legal reasoning.

By Jaehoon Jeong, Jay-Yoon Lee
arXiv AI
Jul 3

AI Assistance for Human Review of Default Judgments

arXiv:2607. 01256v1 Announce Type: cross Abstract: Overwhelmed courts in the United States review millions of default judgments each year.

By Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal, Aviv Caspi, Tatsunori Hashimoto, Carlos Guestrin, David Freeman Engstrom
arXiv AI
Aug 25

Why we need an AI-resilient society- Profiling Large Language Models

The article discusses the evolution of AI across three generations—from explicit logic to neural networks to large language models (LLMs)—and how LLMs introduce new systemic risks. It applies a forensic‑psychology profiling method to identify ten key features of LLMs, such as hallucinations, bias, and cognitive atrophy, revealing an entity that confabulates, amplifies user biases, and erodes human competence. The report concludes with a four‑pillar framework for AI resilience, emphasizing cognitive sovereignty, measurable control, partial autonomy, and openness to safeguard society.

By Thomas Bartz-Beielstein, Eva Bartz