Civil litigation is inherently a life-cycle process: what a lawyer drafts on day one constrains what unfolds at trial months later. Yet existing legal benchmarks evaluate isolated subtasks, and prior legal-agent simulators reinitialize each scenario from shared ground truth, leaving cross-stage causal dependencies unmodeled.
The paper introduces OBJECTION, an inference-time pipeline that adds an Adversarial Lawyer Agent to each of the three reasoning steps—offense, unlawfulness, and culpability—in legal judgment prediction models. By actively injecting defense arguments, the agent challenges the model’s default assumption of guilt, which is common in datasets biased toward guilty outcomes. Using a new Natural Innocent dataset of 3.4k real cases, OBJECTION reduces the False Guilty Rate from 82.93% to 16.69%, demonstrating significant improvement in substantive legal reasoning.
By Jaehoon Jeong, Jay-Yoon Lee
The paper investigates how a defendant’s courtroom statement influences decisions made by large language model (LLM)–simulated jurors. Using the newly introduced JuryBench benchmark, the authors analyze 432,000 verdicts from 20 frontier LLMs, varying defendant background, emotional appeal, and rebuttal content. Results show that emotional persuasion can backfire, that jurors are harsher toward defendants from different backgrounds and more lenient toward same‑background defendants, and that juror ideology strongly shapes verdict severity.
By Cho-Ying Wu
LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement...
The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.
By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
By Jeffrey Flynt