arXiv AI By Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe, Desmond Ong, Dan Jurafsky, Diyi Yang

Warning labels shift perceptions of sycophantic AI, but not its influence

Read the original on arXiv AI →

arXiv:2606. 21317v2 Announce Type: replace-cross Abstract: Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

The paper introduces Safety Nudges, a browser-based tool that displays lightweight, in situ flags when a conversational AI exhibits risky behavior such as hallucination or overconfidence. In a two‑week field study with 45 frequent chatbot users, participants reported that the nudges were useful, clear, and minimally disruptive, and most felt more aware of potential AI harms. However, increased awareness did not automatically translate into measurable changes in user behavior, underscoring the need for relevance, calibration, and user control in nudge design.

By Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
arXiv AI
Aug 25

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

arXiv:2608.21841v1 Announce Type: new Abstract: Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI...

By Rachel Poonsiriwong (Pub), Chayapatr (Pub), Archiwaranguprok, Constanze Albrecht, Monchai Lertsutthiwong, Pattie Maes, Pat Pataranutaporn
arXiv AI
Aug 26

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

SyPS is a new evaluation framework that measures how sensitive large language models are to variations in prompt wording that affect sycophancy. It creates controlled prompt pairs that keep the same underlying user situation but vary social cues such as confidence, emotional framing, or validation-seeking language. The framework introduces the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level metric that separates baseline sycophancy from prompt-induced shifts, allowing model-level comparisons of robustness to social cues.

By Lijia Huang, Yao Fu, Sihao Ren