arXiv:2607. 01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2501. 14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption.
By Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such
arXiv:2607. 01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa
SSP-Bench is a dynamic benchmarking framework designed to evaluate large language models on safety, security, and privacy (SSP) by generating evaluation instances on demand while maintaining domain consistency. It ensures label validity through external sources, enforces scope with service-specific validation, and calibrates difficulty using a multi-model steering panel, framing benchmark construction as a multi-objective optimization over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench exposes systematic failures of static benchmarks, such as near-zero correlation in safety rankings, strong safety–over-refusal coupling, and hidden within-family regressions.
By Fatih Deniz, Yazan Boshmaf, Issa Khalil