arXiv AI By Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma

Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

Read the original on arXiv AI →

arXiv:2608. 13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction

The paper introduces OBJECTION, an inference-time pipeline that adds an Adversarial Lawyer Agent to each of the three reasoning steps—offense, unlawfulness, and culpability—in legal judgment prediction models. By actively injecting defense arguments, the agent challenges the model’s default assumption of guilt, which is common in datasets biased toward guilty outcomes. Using a new Natural Innocent dataset of 3.4k real cases, OBJECTION reduces the False Guilty Rate from 82.93% to 16.69%, demonstrating significant improvement in substantive legal reasoning.

By Jaehoon Jeong, Jay-Yoon Lee
arXiv Computation and Language
Sep 10

When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors

The paper investigates how a defendant’s courtroom statement influences decisions made by large language model (LLM)–simulated jurors. Using the newly introduced JuryBench benchmark, the authors analyze 432,000 verdicts from 20 frontier LLMs, varying defendant background, emotional appeal, and rebuttal content. Results show that emotional persuasion can backfire, that jurors are harsher toward defendants from different backgrounds and more lenient toward same‑background defendants, and that juror ideology strongly shapes verdict severity.

By Cho-Ying Wu
arXiv AI
Sep 10

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

By Nirav Patel, Emily Wenger, Christopher Buccafusco