arXiv AI By Hobin Kim, Xiaoyuan Wu, Omer Akgul, Lujo Bauer, Nicolas Christin

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

Read the original on arXiv AI →

arXiv:2606. 18062v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used to fulfill users' information needs; users ask LLMs about the weather, pose educational questions, and consult them for legal assistance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

The study investigates how users perceive the helpfulness and privacy-preservation of large language model (LLM) responses in privacy-sensitive scenarios. Using 94 participants and 90 PrivacyLens scenarios, researchers found that users’ evaluations of identical LLM outputs varied widely, whereas five proxy LLM judges showed high agreement but low correlation with user judgments. The results suggest that proxy LLMs cannot reliably estimate users’ diverse perceptions of utility and privacy, highlighting the need for more user-centered evaluation methods.

By Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue
arXiv AI
Aug 19

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

The paper proposes a methodological framework to assess large language models (LLMs) as surrogate experts in security surveys, particularly for Security Operations Centres (SOCs). By comparing persona-based and aggregate LLM-generated responses to real SOC professional data, the study evaluates stability, inter-model agreement, and alignment with human answers. Findings reveal that while LLMs produce internally consistent responses, they systematically diverge from experts, showing reduced variance, central tendency bias, and homogenised opinions, indicating they are suitable for piloting and hypothesis generation but not for replacing expert elicitation.

By Despoina Giarimpampa, Roland Meier, Tegawend\'e F. Bissyand\'e, Vincent Lenders, Jacques Klein
arXiv AI
Jun 17

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

arXiv:2604. 09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasingly marketing autonomous AI solutions for SOCs.

By Souradip Nath, Chih-Yi Huang, Aditi Ganapathi, Kashyap Thimmaraju, Jaron Mink, Gail-Joon Ahn
arXiv Computation and Language
Sep 1

WildSEEK: Evaluating Language Models for Information-Seeking

WildSEEK is a new dataset of 3,000 real user information‑seeking queries, manually annotated for risk‑sensitive domains and whether the query is factoid or analytical. The accompanying evaluation framework tests LLM responses against four failure criteria—sycophantic behavior, overreliance, a default US‑centric perspective, and poor handling of vulnerable populations—finding higher failure rates for analytical queries. The authors also train classifiers on WildSEEK to analyze over 1.8 million realistic queries, revealing that more than a third are high‑risk and often analytical.

By Tanise Ceron, Joachim Baumann, Elisa Bassignana, Berat Cabuk, Dirk Hovy, Debora Nozza