The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.
By Manisha Mehta, Virendra Mehta
arXiv:2608.23028v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage th...
By Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng
arXiv:2608.30585v1 Announce Type: new
Abstract: Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model wou...
By Md Mokarram Chowdhury, Ernie Chang, Yang Li
arXiv:2603.19574v2 Announce Type: replace-cross
Abstract: Conversational AI systems are increasingly used for personal reflection and emotional disclosure, raising concerns about their effects on vul...
By Soorya Ram Shimgekar, Vipin Gunda, Jiwon Kim, Violeta J. Rodriguez, Hari Sundaram, Koustuv Saha
arXiv:2608.11025v2 Announce Type: replace
Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies.