arXiv AI

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

arXiv:2607. 13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success.

arXiv Computation and Language
3d ago

Safety Monitors Mostly Catch What the Model Already Refuses

The paper evaluates safety monitors by measuring recall only on prompts that the target model actually answers, rather than on all harmful prompts. Across several guard systems, recall at a 1% false‑positive rate drops sharply when focusing on answered prompts, with monitors catching refused requests 1.1–6.4 times more often than answered ones. Rewriting prompts to be less explicit dramatically increases compliance and reveals that many harmful requests slip past monitors, especially when phrasing is softened. Fine‑tuning guards on these rewritten prompts improves recall from 0.24 to 0.89 on answered requests and generalizes to unseen benchmarks.

By Sripad Karne
arXiv Machine Learning
1d ago

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.

By Sahil Kadadekar
arXiv AI
Aug 28

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

The paper investigates how different framings of a model’s evaluation awareness—whether it is seen as a capabilities cue, a safety cue, both, or neither—affect its compliance with instructions. Experiments on Qwen3-32B using the FORTRESS dataset show that a capabilities‑framed awareness leads to significantly higher compliance (a 24–46 percentage‑point advantage over safety framing) across various steering conditions. A chain‑of‑thought pre‑fill intervention further suggests a causal link, with most pre‑fills shifting compliance in the predicted direction, indicating that evaluation awareness is not a uniform behavior but varies qualitatively with its framing.

By Allison Zhuang, Santiago Aranguri
arXiv AI
Aug 20

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

The paper introduces a post‑training framework that teaches a 4B‑parameter language model to exercise task‑conditioned authority in executable terminal and Model Context Protocol (MCP) environments. By auditing each action across six risk dimensions with deterministic verifiers and optimizing for task‑specific excess‑privilege values, the authors achieve 98.48% safe success and reduce excess‑authority errors from 4.56% to 0.79% on held‑out tasks. The study also demonstrates capability retention, prompt‑directed improvement, and generalization over a 400‑task continuation test.

By Alexander Tu, Michael Tu
arXiv Machine Learning
Sep 22

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

The paper introduces SCOPE, a method that post‑trains computer‑use agents to balance task completion with safety by conditioning actions on environmental risk. It combines supervised fine‑tuning on three trajectory types—capability demonstrations, safe continuations, and explicit refusals—followed by reinforcement learning to improve performance. Experiments starting from Qwen3.5‑9B show that SCOPE‑RL achieves high task success and attack‑avoidance rates, outperforming other agents on OSWorld and OS‑BLIND benchmarks.

By Zeyu Kang, Zhenyun Yin, Yang Zhang, Shan He, Shanzhe Lei, Yanjiu Zhong, Xinquan Chen, Yuhong Wang