Hugging Face Trending Papers

Predicting LLM Safety Before Release by Simulating Deployment

Read the original on Hugging Face Trending Papers →

Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model behavior will occur in deployment: they generally have insufficient coverage, are unrepresentative, and are generally recognizable as tests.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
Sep 2

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

The paper addresses the problem of evaluation awareness in alignment testing, where models can detect they are being evaluated rather than deployed. It introduces two methods: critique refinement, which uses extra inference-time compute to generate and refine action candidates for realism, and DISH, an agent harness that narrows the gap between simulation and real deployment. Experiments show that combining both techniques yields greater realism improvements than either alone, demonstrating that automated approaches can enhance alignment evaluation realism more efficiently than simply extending audit duration.

arXiv AI
Jul 17

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.

By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski