arXiv AI By Nada Hanad, Mehdi Acheli, Ali NourEldin, Mohamed Sellami, Walid Gaaloul

Reliable Remediation Impact Prediction for Black-Box Security Ratings

Read the original on arXiv AI →

arXiv:2607. 16357v1 Announce Type: cross Abstract: Security rating platforms summarize externally observable cyber exposure and are expected to help organizations prioritize remediation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.

By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)
arXiv AI
Jun 6

Explainable AI-Driven Cyber Risk Analytics and Model Reliability Assessment for Intelligent Governance of U.S. Critical Infrastructure: An XGBoost and SHAP-Based Intrusion Detection Framework

arXiv:2606. 05710v1 Announce Type: cross Abstract: The increasing penetrations of the critical infrastructure sector in the United States with intelligent digital technologies have greatly increased exposure to advanced cyber adversaries and operational vulnerabilities.

By B. M. Taslimul Haque, Md. Arifur Rahman, Md. Serajul Kabir Chowdhury Rubel, Md. Iqbal Hossan
arXiv AI
Aug 24

Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

The paper introduces a diagnostic framework for long‑horizon security LLM agents that uses checkpoints to distinguish failures occurring before and after a model’s capability is exposed, and applies controlled interventions to pinpoint upstream bottlenecks. The methodology is tested on four task families—delayed reuse of discovered information, reuse of observed state, recovery from failed strategies, and decision making after uncertain outcomes—revealing that many failures happen before the agent observes the state it later needs to reuse. Experiments with Gemini 2.5 Flash and Gemini 3.7 Flash show that targeted protocol‑disambiguation guidance can significantly alter state observation rates and that the primary source of failure can shift across model generations, underscoring the need for fine‑grained failure diagnostics rather than relying solely on overall task success.

By Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
arXiv AI
Jul 31

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

arXiv:2607. 26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities.

By Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li