arXiv AI By Rose Cymbler, Daniel Guez, Laurent Fabre

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

Read the original on arXiv AI →

arXiv:2608. 09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

arXiv:2608. 08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates.

By Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv Computation and Language
Sep 1

Annotated Surrogate Retrieval for Polish Statutory Law

The paper introduces three retrieval methods for Polish statutory law that use language‑model annotations attached to articles as surrogates. The methods—ASCR, ASCR‑H, and DTF—vary in cost and quality, with ASCR‑H achieving the highest rank‑one accuracy on bar exam questions, while DTF offers competitive performance with lower latency and cost. Extensive evaluation against 14 baselines on 300 exam questions demonstrates significant improvements in head‑rank accuracy and discusses limitations such as coverage asymmetry and negative results for lemmatisation, pseudo‑relevance feedback, and query rewriting.

By Orkun Yi\u{g}it Cengiz
arXiv Computation and Language
6d ago

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

Large language models (LLMs) are increasingly used for legal research, but their fixed training cutoffs and reliance on static knowledge clash with the evolving nature of statutory law. This study introduces a benchmark of 312 expert‑validated, time‑sensitive German statutory QA pairs that examine two temporal failure modes: post‑cutoff staleness and recency bias. Five LLMs were evaluated under four inference settings, and the results show that retrieval‑augmented approaches that enforce temporal validity significantly improve performance, while web search yields unstable gains and a pronounced recency bias.

By Max Prior, Andreas Schultz, Matthias Grabmair