arXiv:2607. 07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model.
By Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa
arXiv:2607. 14285v1 Announce Type: cross Abstract: Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict?
By Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
The paper addresses the problem of evaluation awareness in alignment testing, where models can detect they are being evaluated rather than deployed. It introduces two methods: critique refinement, which uses extra inference-time compute to generate and refine action candidates for realism, and DISH, an agent harness that narrows the gap between simulation and real deployment. Experiments show that combining both techniques yields greater realism improvements than either alone, demonstrating that automated approaches can enhance alignment evaluation realism more efficiently than simply extending audit duration.
arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.
By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski