arXiv AI

Predicting LLM Safety Before Release by Simulating Deployment

arXiv:2607. 07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model.

arXiv AI
Sep 3

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

The paper introduces two methods to make alignment evaluations more realistic: critique refinement, which adds inference-time compute to generate and refine candidate actions, and DISH, a deployment-imitating harness that narrows the gap between simulation and real deployment. Experiments on multiple target models show that combining both techniques yields greater realism improvements than using either alone. The study demonstrates that automated approaches can enhance evaluation realism more efficiently than simply extending audit duration.

By Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes
Hugging Face Trending Papers
Sep 2

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

The paper addresses the problem of evaluation awareness in alignment testing, where models can detect they are being evaluated rather than deployed. It introduces two methods: critique refinement, which uses extra inference-time compute to generate and refine action candidates for realism, and DISH, an agent harness that narrows the gap between simulation and real deployment. Experiments show that combining both techniques yields greater realism improvements than either alone, demonstrating that automated approaches can enhance alignment evaluation realism more efficiently than simply extending audit duration.

arXiv AI
Jul 17

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.

By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski
arXiv AI
Jun 17

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

arXiv:2606. 17114v1 Announce Type: cross Abstract: AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information.

By Hankyul Baek, Jaewon Noh, Sang Seo, Yongsu Kim, Gabriel Waikin Loh Matienzo, Young Il Kim, Ee Wei Seah, Akriti Vij
arXiv AI
Sep 3

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.

By Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk