arXiv AI By Zhengkai Tu, Mingda Zhang, Zijia Wang, Xiaoying Tang, Jimmy Huang

JusticeAxis: Benchmarking Legal Judgment between Rigid Rule Application and Ungrounded Discretion

Read the original on arXiv AI →

JusticeAxis is a new benchmark that formalizes legal judgment as a reference‑anchored task, requiring a single decision tied simultaneously to the statute and the surrounding circumstances. It comprises 256 real‑world criminal cases from 18 countries, each with audio, image, and text evidence, and three lawyer‑written judgments per case: the recorded judgment and two failure judgments. The authors also introduce JusticeAgent, a modular system that first establishes facts and then applies the law, with skills distilled from execution trajectories and validated under Bayesian credible bounds. Experiments reveal that open‑weight backbones tend to drift toward unsupported grounds while frontier models default to statutory interpretations, and that JusticeAgent can effectively plug into commercial systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv AI
Jul 3

AI Assistance for Human Review of Default Judgments

arXiv:2607. 01256v1 Announce Type: cross Abstract: Overwhelmed courts in the United States review millions of default judgments each year.

By Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal, Aviv Caspi, Tatsunori Hashimoto, Carlos Guestrin, David Freeman Engstrom
arXiv AI
Aug 26

Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks

The paper introduces a dual‑judge evaluation protocol for vision‑language models in legally grounded tasks, pairing a 0‑10 quality judge with a strict binary semantic‑equivalence judge. Using a controlled UK traffic‑sign interpretation task, the authors analyze 4,680 evaluations across visibility and occlusion conditions, finding moderate association between judges and an asymmetric Type II error pattern that is most pronounced under heavy occlusion. The protocol requires only one additional LLM call and reveals quality‑trustworthiness signals that single‑judge methods miss.

By Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh