arXiv AI

Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level

arXiv:2608. 14692v1 Announce Type: cross Abstract: Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior.

arXiv AI
Jun 9

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

arXiv:2606. 09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories.

By Yanyan Luo, Xue Han, Ruiqiao Bai, Xin Huang, Yitong Wang, Qian Hu, Qing Wang, Chunxu Zhao, Jie Liu, Cong Geng, Lehao Xing, Pengwei Hu, Junlan Feng
arXiv AI
Sep 12

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms

The study audits Bluesky’s Moderation Service (BMS) using its 10.6 million public moderation labels from 2025. It finds that BMS operates as a human‑AI collaboration: sexual and graphic content is flagged automatically in seconds, while more nuanced or high‑stakes content requires human review that can take hours or days. The system shows high precision (0.837) but low recall (0.222), with annotators detecting 4.5 times more harmful content than the system, and clustering reveals harms ranging from hostility toward protected groups to the spread of explicit material.

By Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar, Vabuk Pahari, Abhijnan Chakraborty, Krishna P. Gummadi, Ingmar Weber, Abhisek Dash
arXiv Machine Learning
2d ago

Privacy in Personalized AI Is a System Property, Not Just a Model Property

The paper argues that privacy in personalized AI should be viewed as a system-level issue rather than just a model-level one. It identifies four interconnected privacy‑risk channels in personalized AI and proposes four system‑level requirements—interaction trajectories, internal information flows, indirect leakage, and the privacy‑utility trade‑off—for evaluating privacy. The authors call for these requirements to be systematically incorporated into privacy audits of personalized AI systems.

By Guillaume Salha-Galvan, Jiaying Xu
Hugging Face Trending Papers
Aug 11

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. However, the 'black-box' nature of these models, combined with the high dimensionality of multimodal cyber-physical data, poses significant interpretability challenges.

arXiv AI
Sep 24

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

The paper introduces Safety Nudges, a browser-based tool that displays lightweight, in situ flags when a conversational AI exhibits risky behavior such as hallucination or overconfidence. In a two‑week field study with 45 frequent chatbot users, participants reported that the nudges were useful, clear, and minimally disruptive, and most felt more aware of potential AI harms. However, increased awareness did not automatically translate into measurable changes in user behavior, underscoring the need for relevance, calibration, and user control in nudge design.

By Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
arXiv AI
Sep 10

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

PersonaTeaming introduces a workflow that incorporates personas into adversarial prompt generation for generative AI, achieving higher attack success rates than the state‑of‑the‑art RainbowPlus while preserving prompt diversity. The system is extended into a user‑facing playground that lets red‑teamers create their own personas and collaborate with AI to refine prompts, fostering diverse strategies. A user study with 11 industry practitioners found the playground produced useful outputs and encouraged out‑of‑the‑box thinking, even when suggestions were not strictly followed.

By Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim, Akshita Jha, Lauren Wilcox, Kenneth Holstein, Motahhare Eslami, Leon A. Gatys
arXiv AI
Sep 24

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

The paper introduces the Systemic Risk Index, an open pipeline and dashboard that aggregates evidence from 19 public AI benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice. It evaluates 18 models using harm‑preserving perturbations and simulated deployment contexts, offering users the ability to switch between average and worst‑case aggregation and to trace each risk rating back to its benchmark evidence. The study finds that worst‑case scores can be 14 to 37 points lower than average scores, and that LLM judges agree with human graders at a level comparable to human‑human agreement.

By Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin
arXiv AI
Jun 2

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

arXiv:2606. 00033v1 Announce Type: cross Abstract: While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments.

By Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi, Philip Quirke, Austin Meek, Fazl Barez, Amirali Abdullah