arXiv AI

Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations

The article evaluates the Sola Security Brain, a purpose-built security intelligence layer, against a general-purpose coding agent (Claude Code) on 28 cloud‑security investigation tasks. The Sola Security Brain achieved 0.693 coverage versus 0.387 for the coding agent, a 79.2% relative gain, and outperformed the agent on 25 of 28 tasks while incurring far lower reasoning and cost per unit of coverage. The study also identifies a ‘sample‑and‑generalise’ pattern where the live agent reports universal negatives based on limited sampling, illustrating a potential efficiency trade‑off in cloud investigations.

arXiv AI
Aug 28

BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click

BekchiAI introduces a benchmark and platform for evaluating large language model agents. The benchmark comprises 13 tool‑using ReAct agents across seven task categories, totaling 2,057 deterministic test tasks with verifier‑checkable gold answers. The platform offers web‑based observability, token and latency telemetry, and remote run termination for live agents.

By Mesut Toruk
arXiv AI
Jul 31

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.

By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
arXiv Machine Learning
Sep 3

Context Inference Attacks Without Jailbreaks

The paper investigates privacy risks in agentic AI systems that assemble sensitive data into a hidden context before responding. It introduces context‑inference attacks, a security game that evaluates how well attackers can recover this hidden context under varying levels of knowledge and indirect delivery. Experiments show that even with controls such as instructions not to disclose, logit suppression, and context dilution, agents can leak significant contextual information, achieving high success rates across multiple attack settings.

By Prince Jha, Samuele Poppi, Nils Lukas
arXiv AI
Sep 10

Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain

The study investigates whether AI coding assistants check trust signals before installing software. Researchers pre‑registered a controlled experiment on six open‑source research projects, creating nine modified versions per project with varying trust signals and running 1,920 trials across three models and two operating modes. Results showed that verification of trust signals was almost nonexistent—only 0.5% of trials involved any signal inspection, and no trial executed a verification command, indicating that publishing signals alone does not ensure secure behavior.

By Pengyin Shan
arXiv AI
Sep 4

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

SENTINEL‑RL is an agentic SOC architecture that separates topological reasoning from semantic reasoning. It uses a heterogeneous graph attention encoder to compress a large authentication subgraph into a fixed‑dimensional state, a PPO policy to map that state to constrained investigative actions, and an LLM loop that only consumes policy recommendations and produces analyst‑readable narratives. Experiments on LANL and Indiana University datasets show fast graph ingestion, reliable alerting, high PPO performance, and a median 6.3‑second containment cycle.

By Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild
arXiv AI
Aug 24

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

ClawSentry is an open‑source, framework‑agnostic security supervision gateway designed to protect autonomous large language model (LLM) agents from progressive risks that can arise at four points in the agent control loop: skill admission, invocation‑time intent, execution‑time effect, and post‑action consequence. It introduces a multi‑tier decision engine—deterministic L1, rule‑anchored L2, and read‑only L3—alongside a First‑Use Skill Package Review (FSPR) and an Agent Harness Protocol (AHP) that applies a single policy across multiple agent runtimes without modifying their internals. Evaluation on SkillInject and the SkillsSafety benchmark shows that ClawSentry significantly reduces contextual adversarial skill risk (ASR) while maintaining high task success rates (TSR).

By Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu