Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark
Read the original on arXiv Computation and Language →The paper investigates whether fine‑tuning improves how language models use supplied Bangladeshi legal text in bilingual question‑answering. Using a hierarchy‑preserving statutory corpus, 2,165 fine‑tuning examples, and a 150‑item control set, the authors evaluate six instruction‑tuned models with multiple LoRA seeds, separating scoring, retrieval, and model effects. Results show that while fine‑tuning can boost overall accuracy, it does not increase the models’ reliance on the governing provision, highlighting the need to disentangle scorer, retriever, and model contributions in legal adaptation studies.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.