arXiv Machine Learning By Albert Sawczyn, Jakub Binkowski, Kamil Tagowski, {\L}ukasz Augustyniak, Berenika Kaczmarek-Templin, Tomasz Kajdanowicz

Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv Computation and Language
Sep 24

LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning

LEGO is a dual‑module framework that combines a Legal Expert GraphRAG system with an expert Chain‑of‑Thought approach to enhance complex legal reasoning. The GraphRAG component uses an expert‑annotated civil code graph and a greedy normative‑coverage retrieval algorithm to extract relevant provision subgraphs, while the Chain‑of‑Thought module structures retrieved provisions and case facts into a Provision‑Fact‑Conclusion reasoning flow. Using a Qwen3‑8B backbone, LEGO achieves 40.53% exact‑match accuracy on LawExamQA_Civil, surpassing baseline RAG and CoT models and matching larger models on multi‑hop and open‑ended benchmarks, with ablation studies confirming the complementary benefits of both modules.

By Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng, Yukun Yan, Zhi Zheng, Antonino Rotolo, Yun Liu, Weixing Shen
arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu