arXiv AI By Veronika Batzdorfer (KIT), Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\'edialab, Sciences Po)

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

Read the original on arXiv AI →

The study shows that while large language model (LLM) annotations of stakeholder consultation submissions are highly reproducible (intraclass correlations > 0.99), they do not reliably capture the intended construct measured by structured survey responses. Divergence between LLM-inferred and survey measures varies by stakeholder group, with business associations expressing more AI risk concern in text than in surveys, and spatial autocorrelation indicates neighboring European countries share similar text-based stances. Despite these divergences, survey-reported concerns remain strongly linked to support for explainability across all levels of divergence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

The paper argues that modern inference pipelines add an unseen layer of control between a language model’s frozen weights and its output, altering probability distributions before token selection. It introduces the concepts of the Inference Attribution Problem, Probability Placement, and Inference Policy Transparency to describe how such interventions can bias generated language toward specific frames and how these biases cannot be traced solely to model weights. The authors discuss the governance, security, and economic implications of these undisclosed inference policies, referencing EU AI Act, Digital Services Act, and FTC doctrines.

By Augusto Camargo
arXiv AI
Aug 19

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

The paper proposes a methodological framework to assess large language models (LLMs) as surrogate experts in security surveys, particularly for Security Operations Centres (SOCs). By comparing persona-based and aggregate LLM-generated responses to real SOC professional data, the study evaluates stability, inter-model agreement, and alignment with human answers. Findings reveal that while LLMs produce internally consistent responses, they systematically diverge from experts, showing reduced variance, central tendency bias, and homogenised opinions, indicating they are suitable for piloting and hypothesis generation but not for replacing expert elicitation.

By Despoina Giarimpampa, Roland Meier, Tegawend\'e F. Bissyand\'e, Vincent Lenders, Jacques Klein