arXiv Computation and Language

Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark

The paper investigates whether fine‑tuning improves how language models use supplied Bangladeshi legal text in bilingual question‑answering. Using a hierarchy‑preserving statutory corpus, 2,165 fine‑tuning examples, and a 150‑item control set, the authors evaluate six instruction‑tuned models with multiple LoRA seeds, separating scoring, retrieval, and model effects. Results show that while fine‑tuning can boost overall accuracy, it does not increase the models’ reliance on the governing provision, highlighting the need to disentangle scorer, retriever, and model contributions in legal adaptation studies.

arXiv AI
Jun 18

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv:2606. 18699v1 Announce Type: cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored.

By Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang
arXiv Computation and Language
Sep 1

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

IndicQE-APE is a consolidated benchmark that unifies quality estimation (QE) and automatic post‑editing (APE) data for nine Indic language pairs, comprising 126,754 instances with multiple aligned labels such as direct assessment, human post‑edit, word‑level OK/BAD tags, and error explanations. The dataset includes a stratified test set across four difficulty axes and supports training and evaluation of six prompted large language models, three COMET metrics, and three APE systems. Experiments reveal that segments with conflicting holistic and token‑level quality signals are consistently ranked lower, while annotator disagreement shows no effect when controlled for score distribution. whyItMatters":"The benchmark provides a unified resource for training and evaluating QE and APE across Indic languages, enabling consistent comparison of models and metrics on a shared dataset."

By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv Computation and Language
6d ago

KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records

KhatianDoc is a new benchmark that tests multimodal large language models on Bengali legal land records, specifically the handwritten RS Khatians used in Bangladesh. The benchmark comprises four tasks—symbol recognition, base‑16 to decimal conversion, structured field extraction, and legal document question answering—drawn from 107 real records and 1,634 QA pairs. Six multimodal LLMs were evaluated under a zero‑shot protocol, revealing that many models fail to answer a significant portion of questions correctly and perform poorly on arithmetic conversion, highlighting a lack of capability rather than a performance gap.

By Tasmiad Hasan, Arafat Zaman Ratul, Sarker Sadman Saalim, S. M. Shah Nawaz Hossain, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi
arXiv AI
Aug 19

Grading Needs a Rubric, Not Intelligence

Small language models can grade open‑ended exam answers as reliably as much larger models when they use an explicit rubric. In experiments with six cost‑efficient model configurations, the rubric decouples grading from judge intelligence, with answer identity explaining 95.6% of score variance and judge identity only 0.2%. Removing rubric criteria or the official answer collapses reliability and inflates scores, showing the rubric’s essential role.

By Jhen-Ke Lin
arXiv AI
Aug 20

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

The paper introduces THPT‑Ladder, a benchmark based on Vietnam’s 2025 National High School Graduation Examination’s convex grading scheme, which rewards partial credit non‑additively. It shows that standard accuracy metrics inflate model scores because they treat partial knowledge proportionally, whereas the official rubric penalizes incomplete correct sets. Using the benchmark, the authors demonstrate that this discrepancy can shift a model’s percentile ranking by up to 13 points among over 480,000 candidates.

By Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh