Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.
By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv:2607. 01256v1 Announce Type: cross Abstract: Overwhelmed courts in the United States review millions of default judgments each year.
By Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal, Aviv Caspi, Tatsunori Hashimoto, Carlos Guestrin, David Freeman Engstrom
arXiv:2605. 28183v2 Announce Type: replace-cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law.
By Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Ko\c{c}ak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair
arXiv:2603. 22973v2 Announce Type: replace Abstract: Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity.
By Avrile Floro (UPHF), Tamara Dhorasoo (UPHF), Soline Pellez (UPHF), Nils Holzenberger
The paper introduces a dual‑judge evaluation protocol for vision‑language models in legally grounded tasks, pairing a 0‑10 quality judge with a strict binary semantic‑equivalence judge. Using a controlled UK traffic‑sign interpretation task, the authors analyze 4,680 evaluations across visibility and occlusion conditions, finding moderate association between judges and an asymmetric Type II error pattern that is most pronounced under heavy occlusion. The protocol requires only one additional LLM call and reveals quality‑trustworthiness signals that single‑judge methods miss.
By Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh
arXiv:2608. 12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
The paper introduces the Wiggle Framework, a unified stress test for assessing epistemic stability in large language model (LLM) judges. It evaluates judge robustness across three dimensions—Mechanical Consistency, Single-turn Conviction, and Multi-turn Persistence—using 9 frontier models on 14 judging tasks related to safety, toxicity, AI writing detection, and political-response evaluation. Results show significant instability, with verdict flips ranging from 25–71% under static pushback and 62–91% when challenged by an adversarial LLM, and highlight that successful pressure often misaligns with ground truth.
By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
The paper introduces a risk‑controlled framework for using large language models (LLMs) as judges in tasks without reference answers. By calibrating uncertainty thresholds on a held‑out set, the method ensures that the false discovery rate of accepted verdicts stays below a user‑specified level α with high probability, using finite‑sample Clopper–Pearson intervals. When the parametric judge lacks confidence, the instance is routed to a retrieval‑augmented mode with a second calibrated threshold, preserving the error guarantee while achieving higher coverage than single‑mode baselines.
arXiv:2609.10293v1 Announce Type: new
Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the sou...
By Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos
arXiv:2606. 22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence?
By Jeffrey Flynt
arXiv:2606.09389v2 Announce Type: replace
Abstract: As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses...
By Yifan Chen, Haitao Li, Yiran Hu, Kaisong Song, Jun Lin, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu
arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.
By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan