arXiv Machine Learning By Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi, Omar Ibne Shahid

KrishokChat: A Provenance-Traceable Multi-Task Bengali Agricultural Benchmark with Safety-Critical Chemical Advisory

Read the original on arXiv Machine Learning →

arXiv:2606. 29243v2 Announce Type: replace Abstract: We introduce KrishokChat, an 85,979-instance Bengali agricultural benchmark built from 284 government publications, 13 institutions, and six regional dialects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 25

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

The paper introduces BanglaSafe, a benchmark of 879 Bengali prompts that covers 17 culturally grounded harm categories and five prompting conditions. Evaluation of 18 frontier LLMs shows that 53.6% of responses are unsafe or partially unsafe, with 14.7% containing strictly harmful content. The study finds that the writing style within Bengali has a stronger impact on safety than the language switch itself, and that current safety classifiers struggle to reliably evaluate Bengali content.

By Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady
arXiv AI
Sep 25

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

IndicBankBench is a 799‑case benchmark designed to evaluate the safety and reliability of language model assistants in Indian retail banking. It covers five operational domains, a capability/refusal domain, and twenty primary axes, assessing each case at four stages: safety, action and tool use, response adequacy, and advisory quality. The benchmark uses deterministic safety checks, a narrow resolver for ambiguous confirmation‑before‑write scenarios, and an LLM judge for semantic response adequacy, reporting strict pass rates that reveal a gap between strict reliability (43.7%–58.2%) and at‑least‑once success (60%–74%).

By Suvradip Paul, Chandra Bhushan, Harsh Sharma, Nitin Kukreja, Yatharth Dedhia, Keyur Doshi, Prashant Devadiga