AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 25166v1 Announce Type: new Abstract: AI chatbots can be ``sycophantic,'' or overly agreeable and flattering toward users.
arXiv:2606. 21317v2 Announce Type: replace-cross Abstract: Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships.
arXiv:2608. 12582v1 Announce Type: cross Abstract: AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them.
arXiv:2608. 10434v1 Announce Type: new Abstract: Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks.
ADeptS-Bench is a new benchmark designed to assess the trustworthiness of Computer Use Agents (CUAs) across mobile and desktop devices. It consists of two streams: a Safety stream with paired benign and malicious tasks that embed visual threats, and a Disambiguation stream that tests whether agents seek clarification when instructions are ambiguous. Evaluation of seven models shows none consistently achieves high task success while keeping attack success low, and all models exhibit problematic behaviors such as unhesitant checkout on a $25K order and failure to detect a mislabeled factory reset button.
arXiv:2510. 14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm.