arXiv Computation and Language

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

The paper compares Jev with nine language models on the ContractNLI task, assessing inference cost, response time, average correctness, and correctness under repeated requests. Controlled experiments vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev achieves the lowest cost and median response time, whereas hosted language models show higher baseline accuracy, but rankings differ when evaluating correctness across all conditions and repeats.

arXiv Computation and Language
Aug 31

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

The paper introduces LongJudgeBench, a benchmark designed to evaluate large language models (LLMs) acting as judges for long-form text generation. It highlights that long-form evaluation requires complex, document-level assessments beyond simple length, such as organization, coverage, depth, consistency, and scenario-specific quality. Experiments show a significant reliability gap among current LLM judges, indicating instability across scenarios and limited effectiveness of rubrics or references.

By Junjie Chen, Yuxi Dong, Haitao Li, Weihang Su, Yujia Zhou, Min Zhang, Yiqun Liu, Qingyao Ai
arXiv Computation and Language
Sep 21

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

JudgeSense is a benchmark comprising 880 items from human‑labelled corpora, each presented under two differently worded instructions that ask the same question. The study evaluates 25 judges from six providers across four tasks, measuring how rewording affects agreement with the judge’s own verdicts. Results show that rewording reduces agreement on all tasks, with significant effects on two, and that stability varies across tasks and is not predicted by parameter count.

By Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
arXiv Computation and Language
Sep 1

Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning

The paper reports on building a retrieval‑augmented legal assistant for Uzbek that operates in both a managed cloud service and an on‑premises deployment. It introduces two new domain benchmarks—one for retrieval and one for end‑to‑end QA—and shows that fine‑tuning an open‑weight text embedder (UTE‑1) can close the performance gap with proprietary models under tight cost and latency constraints. The authors also provide negative results for a QLoRA experiment and release the benchmarks, evaluation code, and the fine‑tuned embedder for future low‑resource legal NLP work.

By Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan
arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
Hugging Face Trending Papers
Aug 20

ContractScrub: A benchmark for final review of legal contracts

ContractScrub is a new benchmark that evaluates large language models on the task of contract scrubbing—reviewing legal agreements for errors and inconsistencies. The benchmark includes contracts crafted by experienced lawyers covering diverse error types such as misuse of defined terms, incorrect references, and inconsistent language. Initial tests show that even leading models perform poorly, with only one achieving a 0.75 macro‑average recall, highlighting the gap between general LLM performance and real‑world legal tasks.

arXiv AI
Sep 17

Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

The paper presents a deployed system for answering questions over normative documents that is aware of document version and scope. It evaluates a hosted retrieval service against a governed system that applies explicit rules for version and scope resolution, finding the governed system achieves a higher score (97.7 vs 88.1). The study includes a public benchmark, evaluation scripts, and reports commercial deployment metrics, such as 1,126 users and 100,000 calls per day by April 2026.

By Liuyin Wang, Shuaipeng Jin, Jiwei Shi, Jensen Hsu
arXiv Computation and Language
Sep 1

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.

By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang