arXiv:2607. 21274v1 Announce Type: cross Abstract: We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments.
By Katerina Papantoniou, Panagiotis Papadakos, Theodore Patkos, Dimitris Garefalakis, Nikos Vardakis, Dimitris Plexousakis
arXiv:2605. 21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses.
By Souvick Das, Sallam Abualhaija, Domenico Bianculli
The paper introduces three retrieval methods for Polish statutory law that use language‑model annotations attached to articles as surrogates. The methods—ASCR, ASCR‑H, and DTF—vary in cost and quality, with ASCR‑H achieving the highest rank‑one accuracy on bar exam questions, while DTF offers competitive performance with lower latency and cost. Extensive evaluation against 14 baselines on 300 exam questions demonstrates significant improvements in head‑rank accuracy and discusses limitations such as coverage asymmetry and negative results for lemmatisation, pseudo‑relevance feedback, and query rewriting.
By Orkun Yi\u{g}it Cengiz
arXiv:2608. 09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one.
By Rose Cymbler, Daniel Guez, Laurent Fabre
arXiv:2608. 05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications.
By Ayoub Kirouane, Christos Petrocheilos
The paper presents a cross‑lingual legal QA system for Vietnamese labour law, introducing a bilingual evaluation suite of 231 Vietnamese–English question–answer pairs, 75 of which are annotated for five complex legal reasoning phenomena. It evaluates a verifier‑guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions, and introduces six automatic diagnostics for faithfulness to retrieved evidence. Experiments show that dense retrieval outperforms sparse and hybrid retrieval, translation placement has no significant effect on diagnostics, and verifier‑guided correction modestly improves citation preservation but not other dimensions, with human evaluation indicating a gap between automatic diagnostics and human judgments.
By Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson