arXiv Computation and Language

UK-PRBENCH: A Paragraph-Level Precedent Retrieval Benchmark for United Kingdom Case Law

Hugging Face Trending Papers
Jul 13

Generative Chinese Statute Retrieval

Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries and formal statutory language. In this paper, we propose GCSR, a generative statute retrieval framework that reformulates statute retrieval as a sequence generation problem and internalizes statutory knowledge into a generative model.

arXiv Computation and Language
4d ago

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

The paper introduces an automated pipeline that converts court decisions into legal commentaries for specific German Civil Code sections, using paragraph extraction, summarization, keyword clustering, and large language models to generate headings and citation-rich sections. The system processes 4,555 decisions from the German Federal Court of Justice, evaluates the output on relevance, heading match, citation faithfulness, cluster distinction, and logical ordering, and demonstrates that rapid, low-cost commentary generation is feasible while noting limitations due to source restrictions and legal reasoning norms.

By Max Prior, Niklas Wais, Matthias Grabmair
arXiv AI
Aug 19

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

CoAL‑RAG is a complexity‑aware legal retrieval‑augmented generation method that adapts its retrieval strategy based on a multi‑dimensional evaluation of question essence and retrieval consistency. It quantifies reasoning demand from the logical structure of a question and uses the discrepancy between semantic and keyword retrieval to gauge problem complexity, thereby selecting the most suitable retrieval approach and filtering context dynamically. Experiments show that CoAL‑RAG outperforms baseline models on Chinese legal benchmarks (SocialLawQA, LawBench) with a 42.5% BLEU improvement and 3.6× ROUGE‑L, while also achieving strong cross‑jurisdictional performance on English datasets (LexGLUE, CaseHold).

By Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee
arXiv Computation and Language
Sep 4

LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

LexIssue introduces a benchmark for identifying disputed legal issues in Chinese civil litigation, comprising 430 real‑world cases and 1,303 expert‑annotated issues. The dataset is built around a hierarchical schema that links free‑form issue descriptions to structured legal categories, enabling two complementary tasks: issue generation and issue classification. A retrieval‑augmented knowledge base covering 27 causes of action and 441 issue entries is provided, and experiments show that incorporating this knowledge consistently improves model performance on the tasks.

By Huiyuan Xie, Yuqin Huang, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye