arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
arXiv:2605. 29738v2 Announce Type: replace-cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.
By Volodymyr Ovcharov
arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.
By Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang
arXiv:2609.22529v1 Announce Type: new
Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its t...
By Genis Skura, Roland Bouffanais, Didier Wernli
The paper presents a cross‑lingual legal QA system for Vietnamese labour law, introducing a bilingual evaluation suite of 231 Vietnamese–English question–answer pairs, 75 of which are annotated for five complex legal reasoning phenomena. It evaluates a verifier‑guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions, and introduces six automatic diagnostics for faithfulness to retrieved evidence. Experiments show that dense retrieval outperforms sparse and hybrid retrieval, translation placement has no significant effect on diagnostics, and verifier‑guided correction modestly improves citation preservation but not other dimensions, with human evaluation indicating a gap between automatic diagnostics and human judgments.
By Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson
arXiv:2606. 06679v1 Announce Type: cross Abstract: Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora.
By Xi Xuan, Wenxin Zhang, Yufei Zhou, King-kui Sin, Chunyu Kit
arXiv:2606.09389v2 Announce Type: replace
Abstract: As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses...
By Yifan Chen, Haitao Li, Yiran Hu, Kaisong Song, Jun Lin, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu
The paper surveys how large language models (LLMs) are being applied in legal tasks such as judgement prediction, document analysis, and drafting. It reviews the benefits of automation while highlighting legal challenges like privacy, bias, and explainability. The authors also discuss data resources for legal domain specialization and outline future research directions.
By Zhongxiang Sun
arXiv:2608.28645v1 Announce Type: cross
Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, ther...
By Sindhu Shetty, Spurthi Setty, Natan Vidra
arXiv:2604. 26233v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in particular the factors that lead them to decide difficult questions.
By Oisin Suttle, David Lillis
The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.
By Yao Dou, Benjamin Mamut, Wei Xu
The paper introduces a sentence‑level benchmark for judging large language models’ ability to classify interpretive canons used by the German Federal Constitutional Court, based on Larenz’s framework. It operationalizes these canons as classification criteria, provides a dataset of court decisions annotated at the sentence level, and evaluates four LLMs with both expert hand‑written prompts and prompts optimized via Genetic‑Pareto. The results show mean F1 scores between 70.4 and 79.2, with grammatical interpretation being the easiest and systematic interpretation the hardest, and indicate that expert prompts already offer a strong baseline.
By Felix Ringe