arXiv AI By Suvradip Paul, Chandra Bhushan, Harsh Sharma, Nitin Kukreja, Yatharth Dedhia, Keyur Doshi, Prashant Devadiga

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Read the original on arXiv AI →

IndicBankBench is a 799‑case benchmark designed to evaluate the safety and reliability of language model assistants in Indian retail banking. It covers five operational domains, a capability/refusal domain, and twenty primary axes, assessing each case at four stages: safety, action and tool use, response adequacy, and advisory quality. The benchmark uses deterministic safety checks, a narrow resolver for ambiguous confirmation‑before‑write scenarios, and an LLM judge for semantic response adequacy, reporting strict pass rates that reveal a gap between strict reliability (43.7%–58.2%) and at‑least‑once success (60%–74%).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

FraudBench is a new benchmark that tests policy‑grounded banking conversational agents against adaptive fraud scenarios. It uses a dual‑control framework and a 698‑document internal policy corpus, presenting 150 adversarial scenarios (107 public, 43 held‑out) that require agents to manage mutable account state and tool access while preventing identity, authorization, and trust manipulation. Preliminary results on four agents show attack‑security rates between 49% and 65%, highlighting weaknesses in money‑mule and first‑party fraud detection.

By Dheeraj Mohandas Pai, Lu Xian
arXiv AI
Sep 2

Validity-Aware Jailbreak Evaluation for Large Language Models

The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.

By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran
arXiv Machine Learning
Jun 25

RAS: Measuring LLM Safety Through Refusal Alignment

arXiv:2606. 25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.

By Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee