arXiv Machine Learning

Towards AI epidemiology: a measurement standardisation framework for prospective risk detection

arXiv:2512. 15783v3 Announce Type: replace-cross Abstract: This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for prospective risk detection in deployed AI systems, without access to model internals.

arXiv AI
Sep 18

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

The study shows that while large language model (LLM) annotations of stakeholder consultation submissions are highly reproducible (intraclass correlations > 0.99), they do not reliably capture the intended construct measured by structured survey responses. Divergence between LLM-inferred and survey measures varies by stakeholder group, with business associations expressing more AI risk concern in text than in surveys, and spatial autocorrelation indicates neighboring European countries share similar text-based stances. Despite these divergences, survey-reported concerns remain strongly linked to support for explainability across all levels of divergence.

By Veronika Batzdorfer (KIT), Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\'edialab, Sciences Po)
arXiv AI
Sep 18

Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems

The paper introduces Governance-as-Code (GaC), a framework that translates the EU AI Act’s technical requirements into 43 machine‑checkable acceptance criteria across six compliance modules. GaC runs within a CI/CD pipeline, producing Article‑indexed audit evidence and providing actual Rego policy code. The authors validate GaC on two enterprise deployments, showing it reproduces manual audit findings—including three penalty‑triggering violations—while reducing audit labor by about 75%.

By Rudrendu Kumar Paul, Sourav Nandy
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv AI
Jul 10

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

arXiv:2607. 07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires.

By Gwydion Williams, Sara Zannone, Bilal A Mateen
arXiv AI
Jul 1

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.

By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
Hugging Face Trending Papers
Jul 8

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.

arXiv AI
Sep 24

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.

By Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin