ContractScrub is a new benchmark that evaluates large language models on the task of contract scrubbing—reviewing legal agreements for errors and inconsistencies. The benchmark includes contracts crafted by experienced lawyers covering diverse error types such as misuse of defined terms, incorrect references, and inconsistent language. Initial tests show that even leading models perform poorly, with only one achieving a 0.75 macro‑average recall, highlighting the gap between general LLM performance and real‑world legal tasks.
arXiv:2606. 07904v1 Announce Type: new Abstract: Tool-augmented large language model agents increasingly rely on external APIs, but standard tool schemas describe how to call a tool, not when the tool is causally appropriate or what task state it produces.
By Rahul Suresh Babu, Laxmipriya Ganesh Iyer
The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.
By Yao Dou, Benjamin Mamut, Wei Xu
arXiv:2608. 15857v1 Announce Type: new Abstract: Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management.
By Yishun Wang, Wenjin Yi, Wenkai Li, Zongwei Li, Xiaoqi Li
The paper compares Jev with nine language models on the ContractNLI task, assessing inference cost, response time, average correctness, and correctness under repeated requests. Controlled experiments vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev achieves the lowest cost and median response time, whereas hosted language models show higher baseline accuracy, but rankings differ when evaluating correctness across all conditions and repeats.
By Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He
arXiv:2608. 12342v1 Announce Type: cross Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making.
By Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li