arXiv AI

Scalable Delphi: Large Language Models for Structured Risk Estimation

arXiv AI
Sep 24

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.

By Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin
arXiv Computation and Language
6d ago

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.

By Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia
arXiv AI
Sep 16

Mapping U.S. Federal AI Governance Against Sector Vulnerability

The study evaluates 684 U.S. federal AI governance documents for how they address 14 sectors and 24 AI risks, measuring both breadth and depth of coverage. It finds that risks related to robustness, system security, and governance are more frequently and substantively discussed than socioeconomic, environmental, and emerging risks, and that sectors such as public administration, national security, information, and scientific services receive higher coverage than finance and healthcare. By comparing these coverage patterns with expert vulnerability assessments, the authors identify potential gaps in AI governance that could inform future policy and industry decisions.

By Ho Ting Hung, Angelica Chowdhury, James Teague, Simon Mylius, Spencer Michaels, Peter Slattery, Alexander Saeri, Neil Thompson
arXiv Machine Learning
Jul 28

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

arXiv:2607. 23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer.

By Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
arXiv AI
2d ago

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.

By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)