arXiv AI By David Gringras

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Read the original on arXiv AI →

arXiv:2603. 10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never tested.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.