arXiv AI By Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean

ContractScrub: A benchmark for final review of legal contracts

Read the original on arXiv AI →

arXiv:2608. 20204v1 Announce Type: new Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 20

ContractScrub: A benchmark for final review of legal contracts

ContractScrub is a new benchmark that evaluates large language models on the task of contract scrubbing—reviewing legal agreements for errors and inconsistencies. The benchmark includes contracts crafted by experienced lawyers covering diverse error types such as misuse of defined terms, incorrect references, and inconsistent language. Initial tests show that even leading models perform poorly, with only one achieving a 0.75 macro‑average recall, highlighting the gap between general LLM performance and real‑world legal tasks.

arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu
arXiv Computation and Language
Sep 24

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

The paper compares Jev with nine language models on the ContractNLI task, assessing inference cost, response time, average correctness, and correctness under repeated requests. Controlled experiments vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev achieves the lowest cost and median response time, whereas hosted language models show higher baseline accuracy, but rankings differ when evaluating correctness across all conditions and repeats.

By Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He