arXiv AI By Royce Moon, Lav R. Varshney

Containment Verification: AI Safety Guarantees Independent of Alignment

Read the original on arXiv AI →

arXiv:2605. 09045v2 Announce Type: replace Abstract: Agentic frameworks are the software layer through which AI agents act in the world.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.

By Albert Wu, Nicholas Roberts, Tzu-Heng Huang, Haoran Lin, Gil Friedman, Sungjun Cho, Gabriel Orlanski, Frederic Sala
arXiv AI
2d ago

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Sapien is a policy engine that enforces stateful contextual policies for autonomous AI agents, specifying allowed tool‑call sequences with an extended regular expression that includes stateful predicates, deferred policy generation, and scoped semantic checks. The system maintains performance close to an unconstrained agent while significantly reducing malicious actions, ruling out 93‑95% of attacks on AgentDojo and 62‑85% on Toolathlon, outperforming traditional tool allowlists on long‑horizon tasks.

By Corinn Tiffany, Wen Zhang, Eugene Bagdasarian, Lillian Tsai