FraudBench is a new benchmark that tests policy‑grounded banking conversational agents against adaptive fraud scenarios. It uses a dual‑control framework and a 698‑document internal policy corpus, presenting 150 adversarial scenarios (107 public, 43 held‑out) that require agents to manage mutable account state and tool access while preventing identity, authorization, and trust manipulation. Preliminary results on four agents show attack‑security rates between 49% and 65%, highlighting weaknesses in money‑mule and first‑party fraud detection.
By Dheeraj Mohandas Pai, Lu Xian
The paper introduces SEAV, a verification‑centric framework for evaluating jailbreak attempts against large language models. SEAV decomposes responses into ordered steps and checks both validity and correctness using LLM‑as‑a‑judge and retrieval‑grounded verification. The method reduces false positives by 14.9 percentage points on a strategic‑dishonesty diagnostic and reclassifies 22.1–51.0% of previously successful jailbreaks as invalid across multiple benchmarks.
By Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran
arXiv:2501. 14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption.
By Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such
arXiv:2606. 25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.
By Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee
arXiv:2606. 29243v2 Announce Type: replace Abstract: We introduce KrishokChat, an 85,979-instance Bengali agricultural benchmark built from 284 government publications, 13 institutions, and six regional dialects.
By Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi, Omar Ibne Shahid
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa