arXiv AI

AI Assistance for Human Review of Default Judgments

arXiv:2607. 01256v1 Announce Type: cross Abstract: Overwhelmed courts in the United States review millions of default judgments each year.

arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv AI
Jun 3

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

arXiv:2605. 25240v2 Announce Type: replace-cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs.

By Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan, Sibo Ma, Charles Dickens, Matthew Guillod, Megan Ma, Julian Nyarko
arXiv Computation and Language
Sep 14

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.

By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
arXiv AI
Sep 10

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang