arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
arXiv:2607. 01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2606. 30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure.
By Camilo Chac\'on Sartori
arXiv:2608. 17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.
By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)