The paper investigates whether fine‑tuning improves how language models use supplied Bangladeshi legal text in bilingual question‑answering. Using a hierarchy‑preserving statutory corpus, 2,165 fine‑tuning examples, and a 150‑item control set, the authors evaluate six instruction‑tuned models with multiple LoRA seeds, separating scoring, retrieval, and model effects. Results show that while fine‑tuning can boost overall accuracy, it does not increase the models’ reliance on the governing provision, highlighting the need to disentangle scorer, retriever, and model contributions in legal adaptation studies.
By Moniruzzaman Mahadi, Abrar Mohammed Tanzim Alam, Sayma Siddika Monalisa, Mir Mohammad Asif Abdullah, Swakkhar Shatabda, Md Adnan Arefeen
arXiv:2607. 23446v1 Announce Type: cross Abstract: A small language model can receive the governing statutory provision and still answer incorrectly.
By Moniruzzaman Mahadi, Abrar Mohammed Tanzim Alam, Sayma Siddika Monalisa, Mir Mohammad Asif Abdullah, Swakkhar Shatabda, Md Adnan Arefeen
The paper reports on building a retrieval‑augmented legal assistant for Uzbek that operates in both a managed cloud service and an on‑premises deployment. It introduces two new domain benchmarks—one for retrieval and one for end‑to‑end QA—and shows that fine‑tuning an open‑weight text embedder (UTE‑1) can close the performance gap with proprietary models under tight cost and latency constraints. The authors also provide negative results for a QLoRA experiment and release the benchmarks, evaluation code, and the fine‑tuned embedder for future low‑resource legal NLP work.
By Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan
arXiv:2608. 09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one.
By Rose Cymbler, Daniel Guez, Laurent Fabre
KhatianDoc is a new benchmark that tests multimodal large language models on Bengali legal land records, specifically the handwritten RS Khatians used in Bangladesh. The benchmark comprises four tasks—symbol recognition, base‑16 to decimal conversion, structured field extraction, and legal document question answering—drawn from 107 real records and 1,634 QA pairs. Six multimodal LLMs were evaluated under a zero‑shot protocol, revealing that many models fail to answer a significant portion of questions correctly and perform poorly on arithmetic conversion, highlighting a lack of capability rather than a performance gap.
By Tasmiad Hasan, Arafat Zaman Ratul, Sarker Sadman Saalim, S. M. Shah Nawaz Hossain, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi
The paper presents a retrieval‑augmented generation pipeline for answering regulatory compliance questions in finance. It builds a three‑stage retriever on LegalBERT and a compact 2B–12B generator served with 4‑bit quantization, achieving a Recall@10 of 0.774 on the ObliQA benchmark and improving answer quality via RAFT‑LoRA fine‑tuning. However, the adapted models fail to transfer to Australian case‑law questions, and a closed‑book model performs almost as well while lacking verifiable grounding.
By Tobias Deu{\ss}er, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa