An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.
arXiv:2608. 07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks.
The article discusses how automated red‑teaming can uncover more vulnerabilities at lower cost than human red‑teaming on AI safety benchmarks, yet this comparison conflates measurement with conclusion. It argues that benchmarks only assess harms within a predefined set, leaving a "threat‑model coverage gap" that can hide new risks, as seen in non‑English prompts. The authors suggest that evaluators from deployment contexts distinct from developers are needed to close this gap.
The paper introduces the AI Assessment Sandbox Configurator, an open‑source framework designed to support technical assessment in AI Regulatory Sandboxes (AIRS) mandated by the EU Artificial Intelligence Act. It outlines 11 architectural and governance requirements for infrastructure that enables large‑scale, structured technical testing, and presents a catalogue of tests, a shared data model, dashboards, and reporting tools that harmonise heterogeneous outputs. An early‑stage pilot demonstrated the framework’s harmonisation and reporting capabilities within a live AIRS engagement, contributing to an official Exit Report.
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
arXiv:2607. 07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not.