arXiv AI

Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization

arXiv:2606. 00116v1 Announce Type: cross Abstract: This study introduces a novel architecture of KAN-based BiGRU model for the task of classification and summarization of legal documents in a low-resource multilingual setup.

arXiv Computation and Language
Sep 16

NepKANUN: A RAG-Based Nepali Legal Assistant

The paper introduces NepKANUN, an AI-powered legal assistant designed specifically for Nepali legal texts. Built on a fine‑tuned large language model and integrated into a Retrieval‑Augmented Generation (RAG) framework, it delivers precise answers to natural language legal queries. Evaluated with BERTScore, the system achieved strong F1 scores of 0.82, 0.77, and 0.71 across simple, moderate, and complex question categories, and expert reviews confirm its usability.

By Bhabuk Thapa, Prasiddha Koirala, Ranjit Raut, Sunil Regmi, Bal Krishna Bal
arXiv AI
3d ago

Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights

The paper evaluates legal text classification models for Korean sexual offense cases, comparing traditional machine learning, large language models, and fine‑tuned domain models. Fine‑tuned KLUE‑BERT achieved the highest accuracy of 99.3%, outperforming GPT‑3.5, GPT‑4.0, and other traditional approaches. Explainable AI techniques were used to analyze predictions, revealing linguistic features that influence decisions and highlighting limitations in capturing subtle textual cues, especially in real‑world KICS data.

By Jeongmin Lee
arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu
arXiv Machine Learning
Sep 24

LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies

LexLattice is an extractive summarizer that models a legal act’s hierarchy as a two‑dimensional semantic lattice and consolidates information over it using a masked 2D neural cellular automata before selecting content. The method achieves state‑of‑the‑art ROUGE scores across all 24 languages of EUR‑Lex‑Sum in both multilingual and cross‑lingual settings, outperforming large instruction‑tuned baselines while using only a 1.8 M‑parameter consolidator on a frozen multilingual encoder. A consolidator trained on high‑resource languages transfers almost losslessly to unseen languages, suggesting the model operates on language‑agnostic semantic geometry rather than surface form.

By Sujay Uday Rittikar, Sheela Ramanna
arXiv AI
Sep 3

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

The paper introduces BASSE, a multilingual meta‑evaluation dataset containing 2,040 human‑rated abstractive summaries produced manually or by five LLMs with four prompts. Annotators scored each summary on coherence, consistency, fluency, relevance, and 5W1H using a 5‑point Likert scale. Benchmarking shows proprietary LLM‑judge models best align with human judgments, followed by criteria‑specific automatic metrics, while open‑source judge LLMs perform poorly.

By Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego\~na Altuna
Hugging Face Trending Papers
Jul 21

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the Indian legal context. AILQA leverages a variety of embedding and generative models, including recent Large Language Models (LLMs), to address the unique challenges posed by the intricate and diverse nature of Indian legal texts and to enhance the accuracy and reliability of responses to legal questions.

arXiv Computation and Language
Sep 1

SinLlama -- A Large Language Model for Sinhala

The paper introduces SinLlama, the first decoder‑based open‑source large language model with explicit support for Sinhala. By extending Llama‑3‑8B, adding Sinhala‑specific tokenizer vocabulary, and performing continual pre‑training on a cleaned 10‑million‑token Sinhala corpus, the authors created a model that surpasses both the base and instruction‑fine‑tuned variants of Llama‑3‑8B on three text classification tasks. This work addresses the underrepresentation of low‑resource languages in open‑source LLMs.

By H. W. K. Aravinda, Rashad Sirajudeen, Samith Karunathilake, Nisansa de Silva, Surangika Ranathunga, Rishemjit Kaur