arXiv:2606. 29243v1 Announce Type: new Abstract: We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset for crop advisory in low-resource settings.
By Khan Raiyan Ibne Reza, Omar Ibne Shahid
arXiv:2606. 28992v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, and text generation.
By Zhaoyang Li, Ruijie Zhang, Jiaqi Liu, Zhaoji Sun
The paper introduces BanglaSafe, a benchmark of 879 Bengali prompts that covers 17 culturally grounded harm categories and five prompting conditions. Evaluation of 18 frontier LLMs shows that 53.6% of responses are unsafe or partially unsafe, with 14.7% containing strictly harmful content. The study finds that the writing style within Bengali has a stronger impact on safety than the language switch itself, and that current safety classifiers struggle to reliably evaluate Bengali content.
By Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady
arXiv:2605. 31483v1 Announce Type: cross Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali.
By Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, Md Taukir Azam Chowdhury
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is cri...
IndicBankBench is a 799‑case benchmark designed to evaluate the safety and reliability of language model assistants in Indian retail banking. It covers five operational domains, a capability/refusal domain, and twenty primary axes, assessing each case at four stages: safety, action and tool use, response adequacy, and advisory quality. The benchmark uses deterministic safety checks, a narrow resolver for ambiguous confirmation‑before‑write scenarios, and an LLM judge for semantic response adequacy, reporting strict pass rates that reveal a gap between strict reliability (43.7%–58.2%) and at‑least‑once success (60%–74%).
By Suvradip Paul, Chandra Bhushan, Harsh Sharma, Nitin Kukreja, Yatharth Dedhia, Keyur Doshi, Prashant Devadiga