arXiv AI By Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll

Predicting LLM Safety Before Release by Simulating Deployment

Read the original on arXiv AI →

arXiv:2607. 07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

The paper introduces two methods to make alignment evaluations more realistic: critique refinement, which adds inference-time compute to generate and refine candidate actions, and DISH, a deployment-imitating harness that narrows the gap between simulation and real deployment. Experiments on multiple target models show that combining both techniques yields greater realism improvements than using either alone. The study demonstrates that automated approaches can enhance evaluation realism more efficiently than simply extending audit duration.

By Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes