arXiv AI By Nirav Patel, Emily Wenger, Christopher Buccafusco

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

Read the original on arXiv AI →

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 6

Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools

arXiv:2604. 26233v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions.

By Oisin Suttle, David Lillis
arXiv Computation and Language
Sep 10

When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors

The paper investigates how a defendant’s courtroom statement influences decisions made by large language model (LLM)–simulated jurors. Using the newly introduced JuryBench benchmark, the authors analyze 432,000 verdicts from 20 frontier LLMs, varying defendant background, emotional appeal, and rebuttal content. Results show that emotional persuasion can backfire, that jurors are harsher toward defendants from different backgrounds and more lenient toward same‑background defendants, and that juror ideology strongly shapes verdict severity.

By Cho-Ying Wu