arXiv AI By Yixuan Liu

Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Aug 28

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

The paper evaluates three AI model security scanners—ModelScan, ModelAudit, and Fickling—using a benchmark of 170 Pickle and PyTorch artifacts from 145 families, 135 of which have binary security labels. It distinguishes coverage metrics such as non‑N/A coverage, analysis completion, and definitive security decisions, finding that ModelAudit achieved 100% definitive decisions, Fickling 81.5%, and ModelScan 49.6%. When a definitive judgment was made, ModelScan reached perfect precision, recall, and F1, while Fickling added no unique true positives beyond those found by the other tools.

By Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal
arXiv AI
Sep 25

Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents

The paper proposes using lightweight, calibrated System One decision models—specifically JEV and Laya—to improve autonomous penetration-testing harnesses that rely on large language models (LLMs). It defines four key decision points (finding adjudication, severity recalibration, agent pruning, and confirmation loops) and presents a NeuroSploit case study showing differences in severity distribution, runtime, and grading when using TypeSafe System One. The authors review existing System One specifications, discuss various RL-based training approaches, and introduce Rave, a domain‑adapted model with a proposed training and evaluation framework.

By Joas Antonio dos Santos Barbosa
arXiv AI
Jun 18

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents

arXiv:2606. 18356v1 Announce Type: cross Abstract: Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persistent memory, send messages, modify databases, or trigger harmful code and tool effects.

By Yuchuan Tian, Mengyu Zheng, Haocheng Mei, Ye Yuan, Chao Xu, Xinghao Chen, Hanting Chen, Yu Wang