Models That Know How Evaluations Are Designed Score Safer
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
arXiv:2606. 18129v1 Announce Type: cross Abstract: Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emotionally sensitive interactions over time.
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
arXiv:2607. 24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging.
Anian is a safety‑gated multimodal AI backend designed for perinatal mental‑health support and mindfulness‑intervention routing. It maps user input into a four‑layer hierarchical state representation—emotion, psychosocial constructs, safety risk, and intervention routes—then fuses local and external risk signals to decide whether to generate AI responses or provide fixed safety content. Prototype evaluation on large public corpora showed high classification performance and perfect high‑risk recall in a controlled stress test, though clinical validity remains unestablished.
arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.
arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.
arXiv:2607. 02885v1 Announce Type: cross Abstract: Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction between cognitive and behavioral factors.
arXiv:2607. 25681v1 Announce Type: new Abstract: Cognitive distortion amplifies negative emotions and contributes to mental health disorders.
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
The paper introduces CarryOnBench, an interactive benchmark that tests whether large language models can revise their interpretation of user intent and recover utility while staying safe in multi‑turn conversations. Using 398 harmful‑looking queries with benign intents, the benchmark simulates 5,970 conversations across 14 models, evaluating both intent‑aligned utility and safety with a new metric called Ben‑Util. Results show that models often withhold information due to misinterpretation, but most can recover with clarifications, revealing failure modes such as unsafe and redundant recovery that single‑turn tests miss.
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.