arXiv AI

A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning

The paper introduces a hybrid insider threat detection framework that combines multi-agent simulation, layered SIEM correlation, trust‑adaptive thresholds, behavioral and communication forensics, and Theory‑of‑Mind reasoning. It evaluates four variants—Layered SIEM‑Core, Cognitive‑Enriched SIEM, Evidence‑Gated SIEM, and an Enron‑calibrated version—showing progressively higher actor‑level F1 scores and reduced false positives, especially with evidence gating. Domain‑shift tests reveal that an Enron‑trained email classifier does not transfer to other domains, but fine‑tuning achieves high F1, and scalability tests confirm stable performance up to 1,000 agents.

arXiv AI
Sep 25

On the Effectiveness of Kernel-Level Evidence for Agent Security

The paper introduces the Agent Cross‑Layer Evidence (ACE) corpus, pairing application‑level telemetry with kernel‑level syscall traces to study agent security. It shows that kernel evidence alone is discriminative and that combining it with application‑level data outperforms either layer alone, revealing complementary signals. The study also demonstrates that this cross‑layer approach generalizes to unseen attack families and works across different agent runtimes.

By Spencer King, Zhilu Zhang, Mikhail Kuznetsov, Kay Liu, Baris Coskun, Wei Ding
arXiv AI
Jun 17

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

arXiv:2606. 17555v1 Announce Type: cross Abstract: Banks simultaneously face signature-based fraud (card-not-present attacks, account takeover, ATM cloning) and behavioural financial crime (structuring, layering, mule networks, business email compromise) -- two threat families with fundamentally different detection requirements.

By Joseph Walusimbi, Joshua Benjamin Ssentongo
arXiv Machine Learning
Jul 31

Cybersecurity Detection Classification with Reasoning-enabled Language Models

arXiv:2607. 28460v1 Announce Type: new Abstract: A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day.

By Amol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig, Diana Bolocan, Laura Vasilie, Alexandru Apostu, Chase Helwig, Mihaela Gaman, Michael Brautbar, Edward Raff, Chase Midler, Sven Krasser
arXiv AI
6d ago

Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations

The article evaluates the Sola Security Brain, a purpose-built security intelligence layer, against a general-purpose coding agent (Claude Code) on 28 cloud‑security investigation tasks. The Sola Security Brain achieved 0.693 coverage versus 0.387 for the coding agent, a 79.2% relative gain, and outperformed the agent on 25 of 28 tasks while incurring far lower reasoning and cost per unit of coverage. The study also identifies a ‘sample‑and‑generalise’ pattern where the live agent reports universal negatives based on limited sampling, illustrating a potential efficiency trade‑off in cloud investigations.

By Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok, Konstantin Koutsyi
arXiv AI
Aug 24

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

ClawSentry is an open‑source, framework‑agnostic security supervision gateway designed to protect autonomous large language model (LLM) agents from progressive risks that can arise at four points in the agent control loop: skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence. It introduces a multi‑tier decision engine—deterministic L1, rule‑anchored L2, and read‑only L3—alongside a First‑Use Skill Package Review (FSPR) and an Agent Harness Protocol (AHP) that applies a single policy across multiple agent runtimes without modifying their internals. Evaluation on SkillInject and the SkillsSafety benchmark shows that ClawSentry significantly reduces contextual adversarial skill risk (ASR) while maintaining high task success rates (TSR).

By Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu