arXiv AI By Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

Read the original on arXiv AI →

CoAL‑RAG is a complexity‑aware legal retrieval‑augmented generation method that adapts its retrieval strategy based on a multi‑dimensional evaluation of question essence and retrieval consistency. It quantifies reasoning demand from the logical structure of a question and uses the discrepancy between semantic and keyword retrieval to gauge problem complexity, thereby selecting the most suitable retrieval approach and filtering context dynamically. Experiments show that CoAL‑RAG outperforms baseline models on Chinese legal benchmarks (SocialLawQA, LawBench) with a 42.5% BLEU improvement and 3.6× ROUGE‑L, while also achieving strong cross‑jurisdictional performance on English datasets (LexGLUE, CaseHold).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 13

Generative Chinese Statute Retrieval

Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries and formal statutory language. In this paper, we propose GCSR, a generative statute retrieval framework that reformulates statute retrieval as a sequence generation problem and internalizes statutory knowledge into a generative model.

Hugging Face Trending Papers
Jul 21

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the Indian legal context. AILQA leverages a variety of embedding and generative models, including recent Large Language Models (LLMs), to address the unique challenges posed by the intricate and diverse nature of Indian legal texts and to enhance the accuracy and reliability of responses to legal questions.

arXiv Computation and Language
Sep 4

LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

LexIssue introduces a benchmark for identifying disputed legal issues in Chinese civil litigation, comprising 430 real‑world cases and 1,303 expert‑annotated issues. The dataset is built around a hierarchical schema that links free‑form issue descriptions to structured legal categories, enabling two complementary tasks: issue generation and issue classification. A retrieval‑augmented knowledge base covering 27 causes of action and 441 issue entries is provided, and experiments show that incorporating this knowledge consistently improves model performance on the tasks.

By Huiyuan Xie, Yuqin Huang, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye
arXiv Computation and Language
3d ago

ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations

ViLegalExpert is a large-scale Vietnamese legal benchmark built from real citizen–lawyer consultations, comprising over 172,000 questions across 34 legal domains with professional answers and expert-verified evidence. It supports legal information retrieval, extractive QA, and abstractive QA. Experiments show that while pretrained language models perform well on QA, hybrid retrieval methods achieve the best evidence retrieval, highlighting significant challenges in grounding legal answers to authoritative sources.

By Dat Tien Nguyen, Nghia Hieu Nguyen, Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen