arXiv AI

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

arXiv:2603. 10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never tested.

arXiv AI
2d ago

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.

By Amit Singh Bhatti, Vishal Vaddina
arXiv AI
Sep 3

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

The paper introduces FUSE, a modular framework that evaluates large language models (LLMs) for dangerous capabilities across three orthogonal pipelines: Knowledge (K), Defense (D), and Harm (H). Using a chemical‑biological module, the authors assess 12 commercial LLMs, revealing divergent profiles among models and families, and showing that newer models increase knowledge while only partially improving defense. The framework’s reliability is supported by high cross‑judge consistency and low inter‑pipeline correlations.

By Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang