Hugging Face Trending Papers

ContractScrub: A benchmark for final review of legal contracts

ContractScrub is a new benchmark that evaluates large language models on the task of contract scrubbing—reviewing legal agreements for errors and inconsistencies. The benchmark includes contracts crafted by experienced lawyers covering diverse error types such as misuse of defined terms, incorrect references, and inconsistent language. Initial tests show that even leading models perform poorly, with only one achieving a 0.75 macro‑average recall, highlighting the gap between general LLM performance and real‑world legal tasks.

arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu
Hugging Face Trending Papers
Aug 3

CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMs

Trust is fundamental in modern regulatory ecosystems, and compliance checking plays a critical role in fostering that trust. Regulatory compliance verification is essential for businesses operating in highly controlled environments, as it ensures alignment with sector-specific guidelines across domains such as financial reporting, data privacy, and cybersecurity.

arXiv AI
Jun 12

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.

By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
arXiv Computation and Language
Sep 24

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

The paper compares Jev with nine language models on the ContractNLI task, assessing inference cost, response time, average correctness, and correctness under repeated requests. Controlled experiments vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev achieves the lowest cost and median response time, whereas hosted language models show higher baseline accuracy, but rankings differ when evaluating correctness across all conditions and repeats.

By Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He
arXiv AI
Sep 3

Automated Vulnerability Injection in Smart Contracts Using Large Language Models

The paper presents a method that employs Large Language Models to automatically inject known vulnerabilities into Solidity smart contracts. Using a multi-step validation pipeline, the authors generate nearly 1,000 candidate contracts from real-world sources, ultimately confirming 32 vulnerable variants across 25 vulnerability types. These validated contracts are then used to evaluate the coverage of three static analysis tools, highlighting both complementary strengths and gaps in current detection approaches.

By Luca Migliaccio, Roberto Natella, Naghmeh Ivaki, Nuno Laranjeiro, Marco Vieira
arXiv AI
Sep 11

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

ContractEval is a diagnostic framework that makes active obligations in procedural instructions explicit by representing them as query‑conditioned obligations. It matches these obligations against response or trace evidence, identifying omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output‑contract violations as distinct conformance failures. In tests on audited procedural contracts, ContractEval detects and localizes all injected structural failures that output‑only and trace‑aware LLM judges miss, though it is not a compliance guarantee and remains calibration‑sensitive.

By Praphul Singh, Shanu Kumar, Akshat Agarwal, Ganesh Kumar