CAREBench: A Child-Safety Risk Benchmark for Language Models
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
arXiv:2606. 04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support.
arXiv:2602. 05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support.
arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understandin...
The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.
arXiv:2608. 11200v1 Announce Type: cross Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate.
Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally.
arXiv:2608.21775v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
EvoHarmBench is a dynamic adversarial evaluation framework that simulates how users iteratively modify harmful content to evade moderation. It uses an optimization loop that evolves evasion strategies at the semantic-cluster level while maintaining human readability, and tests 229 semantic sub-clusters across five violation categories derived from 5,002 real-world adversarial samples. The study shows that even state‑of‑the‑art LLM‑based moderators can be bypassed with an 80.3% success rate after twelve iterations, highlighting significant vulnerabilities in current systems.
arXiv:2606. 02423v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve harmful outcomes beyond their capabilities through extended interactions.
arXiv:2604. 17301v2 Announce Type: replace-cross Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances.