arXiv:2608. 08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates.
By Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.
By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv:2505. 02763v2 Announce Type: replace-cross Abstract: One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without calling for much discretion.
By Matthew Dahl, Eric Mart\'inez
The paper introduces three retrieval methods for Polish statutory law that use language‑model annotations attached to articles as surrogates. The methods—ASCR, ASCR‑H, and DTF—vary in cost and quality, with ASCR‑H achieving the highest rank‑one accuracy on bar exam questions, while DTF offers competitive performance with lower latency and cost. Extensive evaluation against 14 baselines on 300 exam questions demonstrates significant improvements in head‑rank accuracy and discusses limitations such as coverage asymmetry and negative results for lemmatisation, pseudo‑relevance feedback, and query rewriting.
By Orkun Yi\u{g}it Cengiz
Large language models (LLMs) are increasingly used for legal research, but their fixed training cutoffs and reliance on static knowledge clash with the evolving nature of statutory law. This study introduces a benchmark of 312 expert‑validated, time‑sensitive German statutory QA pairs that examine two temporal failure modes: post‑cutoff staleness and recency bias. Five LLMs were evaluated under four inference settings, and the results show that retrieval‑augmented approaches that enforce temporal validity significantly improve performance, while web search yields unstable gains and a pronounced recency bias.
By Max Prior, Andreas Schultz, Matthias Grabmair
arXiv:2607. 03325v1 Announce Type: cross Abstract: We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism.
By Giovanni Piccioli, Alessia Fidelangeli, Piera Santin, Pierpaolo Vivo
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
The paper introduces ingest‑time fact compilation, an architecture that preprocesses and compiles corpus data into self‑contained facts with resolved revisions, deletions, and source trust. By storing this compiled state, query‑time models can retrieve answers directly, avoiding costly reconstruction from raw passages. Experiments show that this approach reduces read cost per question by 12.89× and token usage by 21.6× while maintaining accuracy.
By Kyle Wild, Yusuke Takahashi, Asako Uraki
arXiv:2605.25920v2 Announce Type: replace
Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constra...
By Wei Fan, Yining Zhou, Mufan Zhang, Yanbing Weng, Yiran HU, Tianshi Zheng, Baixuan Xu, Chunyang Li, Jianhui Yang, Haoran Li, Yangqiu Song
arXiv:2509. 00761v4 Announce Type: replace Abstract: Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy.
By Boqin Yuan, Ziqi Wang
The paper presents a smartphone‑compatible, retrieval‑augmented language model tailored to Bangladeshi statutory law. By distilling a 9‑billion‑parameter Gemma‑2 teacher into a 2‑billion‑parameter student using supervised fine‑tuning and QLoRA, the authors achieve significant gains in ROUGE‑L and BERTScore on an English benchmark while keeping the model lightweight (1.6 GB) and operable offline on a Pixel 6. The system retrieves from 36,029 statutory passages using a hybrid dense/BM25 approach, and cross‑lingual evaluation shows effective Bangla query handling against an English‑only corpus, with a practicing lawyer rating the responses highly in a pilot study.
By MD. Nafis Kamal, Mahadi Hasan Fahim, Talha Ridwan, Nadifa Zaman, Fariha Roushon Florin, Farig Yousuf Sadeque, Saadat Rafid Ahmed
The paper presents a retrieval‑augmented generation pipeline for answering regulatory compliance questions in finance. It builds a three‑stage retriever on LegalBERT and a compact 2B–12B generator served with 4‑bit quantization, achieving a Recall@10 of 0.774 on the ObliQA benchmark and improving answer quality via RAFT‑LoRA fine‑tuning. However, the adapted models fail to transfer to Australian case‑law questions, and a closed‑book model performs almost as well while lacking verifiable grounding.
By Tobias Deu{\ss}er, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa