arXiv AI By Madhava Gaikwad

A Translational Note on AI Safety Evaluation

Read the original on arXiv AI →

The article discusses how automated red‑teaming can uncover more vulnerabilities at lower cost than human red‑teaming on AI safety benchmarks, yet this comparison conflates measurement with conclusion. It argues that benchmarks only assess harms within a predefined set, leaving a "threat‑model coverage gap" that can hide new risks, as seen in non‑English prompts. The authors suggest that evaluators from deployment contexts distinct from developers are needed to close this gap.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

arXiv:2606. 00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice.

By Andrei Marian Feier, Veysel Kocaman, Yigit Gul, Ahmet Korkmaz, Alexander Thomas, Aleksei Zakharov, Jay Gil, Mehmet Butgul, David Talby
arXiv Machine Learning
Jul 17

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.

By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
arXiv AI
Sep 11

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

The paper introduces a black-box framework for evaluating agentic AI systems, focusing on multi-step vulnerabilities that standard single-turn tests miss. It presents a seven-domain taxonomy linking observable behaviors to risk categories, an automated SAGE-RT red-teaming process generating 120 adversarial scenarios per domain, and a human-validated evaluation using LLM judges. Empirical tests on CrewAI and AutoGen agents show significant governance, privacy, and behavior risks, demonstrating the framework’s ability to uncover critical architectural weaknesses without privileged access.

By Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi
arXiv AI
Sep 24

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.

By Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin