arXiv AI

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

arXiv:2608. 07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations.

arXiv AI
Sep 24

How Children Design and Reason about Trustworthy AI Chatbots

The study explores how children design AI chatbots and what they consider trustworthy. Using a custom chatbot-building environment, 115 learners aged 8‑18 created 119 chatbots and adjusted traits such as confidence, transparency, and formality. Findings show younger children equate trust with purpose‑fulfillment, while older children focus on transparent, calibrated design, and that students calibrate academic chatbots to be more formal and transparent than hobby ones.

By Deniz Ozturk, Jiayu Li, Daksh Pratap Singh, Yasitha Rajapaksha, Fasika Melese, Bahare Riahi, Shiyan Jiang, Qiao Jin, Joey Huang, Veronica Catet\'e, Tiffany Barnes, Xiaoyi Tian
arXiv AI
Sep 24

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

The paper introduces Safety Nudges, a browser-based tool that displays lightweight, in situ flags when a conversational AI exhibits risky behavior such as hallucination or overconfidence. In a two‑week field study with 45 frequent chatbot users, participants reported that the nudges were useful, clear, and minimally disruptive, and most felt more aware of potential AI harms. However, increased awareness did not automatically translate into measurable changes in user behavior, underscoring the need for relevance, calibration, and user control in nudge design.

By Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta
Hugging Face Trending Papers
Jun 29

CAREBench: A Child-Safety Risk Benchmark for Language Models

How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support.

arXiv AI
Sep 10

Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions

arXiv:2601.09721v2 Announce Type: replace-cross Abstract: Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evalua...

By Vahideh Zolfaghari, Leila Mashhadi, Mitra Ahadi, Farzaneh Sedaghatkar, MohammadReza Kargozari