CAREBench: A Child-Safety Risk Benchmark for Language Models
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support.
arXiv:2606. 29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm?
arXiv:2607. 05407v1 Announce Type: cross Abstract: Modern artificial intelligence (AI) systems present profound new risks to child safety.
arXiv:2606. 04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.
The paper introduces KIDBench, a benchmark designed to evaluate the safety of large language models (LLMs) for children aged 7-11. It includes realistic child queries across ten categories, single- and multi-turn prompts, and compares different prompting strategies—no cues, implicit cues, and explicit age instructions—showing that cueing improves safety scores. The study also reveals uneven safety performance across languages and cultures, and presents KIDGuardLlama, a child-safety evaluator, and KIDLlama, a child-safe response model.
arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.
arXiv:2608. 07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations.
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
arXiv:2606. 23884v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions.
arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.
arXiv:2608.21775v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.
arXiv:2608. 14577v1 Announce Type: cross Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis.