arXiv AI By Despoina Giarimpampa, Roland Meier, Tegawend\'e F. Bissyand\'e, Vincent Lenders, Jacques Klein

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

Read the original on arXiv AI →

The paper proposes a methodological framework to assess large language models (LLMs) as surrogate experts in security surveys, particularly for Security Operations Centres (SOCs). By comparing persona-based and aggregate LLM-generated responses to real SOC professional data, the study evaluates stability, inter-model agreement, and alignment with human answers. Findings reveal that while LLMs produce internally consistent responses, they systematically diverge from experts, showing reduced variance, central tendency bias, and homogenised opinions, indicating they are suitable for piloting and hypothesis generation but not for replacing expert elicitation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 17

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

arXiv:2604. 09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasingly marketing autonomous AI solutions for SOCs.

By Souradip Nath, Chih-Yi Huang, Aditi Ganapathi, Kashyap Thimmaraju, Jaron Mink, Gail-Joon Ahn
arXiv Computation and Language
Sep 3

When Persona Attributes Improve Population Alignment in Large Language Models

The paper investigates how persona prompting—using short textual descriptions of individuals—to align large language models (LLMs) with human survey responses. It examines the impact of selecting different persona attributes and finds that not all attribute combinations improve performance, suggesting that the variation in human responses to survey questions may explain mixed results. The study evaluates multiple attribute selection methods across four social surveys, two countries, six LLMs, and twenty prediction tasks, offering guidance on when persona prompting is beneficial and which attribute choices are most effective.

By Leon Fr\"ohling, Jens Rupprecht, Markus Strohmaier, Claudia Wagner
arXiv AI
Sep 18

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

The study shows that while large language model (LLM) annotations of stakeholder consultation submissions are highly reproducible (intraclass correlations > 0.99), they do not reliably capture the intended construct measured by structured survey responses. Divergence between LLM-inferred and survey measures varies by stakeholder group, with business associations expressing more AI risk concern in text than in surveys, and spatial autocorrelation indicates neighboring European countries share similar text-based stances. Despite these divergences, survey-reported concerns remain strongly linked to support for explainability across all levels of divergence.

By Veronika Batzdorfer (KIT), Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\'edialab, Sciences Po)