arXiv AI

Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs

arXiv AI
Jul 16

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

arXiv:2607. 13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it cannot perform, such as contacting emergency services or administering care.

By Eunna Lee, Jungpyo Nam, Sunjun Hwang
arXiv AI
Jul 23

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

arXiv:2607. 19629v1 Announce Type: cross Abstract: Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other.

By Eunna Lee
arXiv AI
Jun 24

One Year Later...The Harms Persist, But So Do We!

arXiv:2606. 23884v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions.

By Annika Marie Schoene, Cansu Canca, Gautham Vijay Kumar, Anson Antony
arXiv AI
Aug 28

A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian

Anian is a safety‑gated multimodal AI backend designed for perinatal mental‑health support and mindfulness‑intervention routing. It maps user input into a four‑layer hierarchical state representation—emotion, psychosocial constructs, safety risk, and intervention routes—then fuses local and external risk signals to decide whether to generate AI responses or provide fixed safety content. Prototype evaluation on large public corpora showed high classification performance and perfect high‑risk recall in a controlled stress test, though clinical validity remains unestablished.

By Lei Wang, Xiao Wang, Lei Li