arXiv:2607. 19292v1 Announce Type: cross Abstract: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios.
By Gjergji Kasneci, Enkelejda Kasneci
The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.
By Samira Maghool, Paolo Ceravolo
The paper proposes a framework that connects structured hazard analysis, component-level testing, and probabilistic system modelling to assess system-level harms from AI in complex sociotechnical systems. It demonstrates the approach using the UK's Real Time Gross Settlement system, showing how adversarial inputs to LLM-based trading can shift AI behaviour, reduce system resilience, and increase the likelihood of cascading bank failures. The framework aims to provide a traceable pathway from model behaviour to systemic outcomes, enabling evidence-based governance of AI in critical infrastructure.
By Paul Vautravers, Oliver Chalkley, Gabriel Downer, Kate S, Damian Ruck
The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.
By Murat Kantarcioglu
arXiv:2607. 24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high.
By Keivan Navaie
arXiv:2607. 29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation.
By Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi, Stefano Silvestri, Francesco Longo, Antonio Puliafito, Giovanni Merlino
arXiv:2607. 01421v1 Announce Type: cross Abstract: Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assurance by test coverage.
By Laxmipriya Ganesh Iyer
arXiv:2607. 23365v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education.
By Muhammad Tukur, Hayatullahi B. Adeyemo, Tao Chen, Nour Ali, Anis Zarrad, Rick Kazman, Marco Agus, Rami Bahsoon
The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.
By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assurance by test coverage. These frameworks implicitly assume deterministic behavior, discrete and auditable change events, and clear component-to-owner mappings.
The paper introduces SCOPE, a method that post‑trains computer‑use agents to balance task completion with safety by conditioning actions on environmental risk. It combines supervised fine‑tuning on three trajectory types—capability demonstrations, safe continuations, and explicit refusals—followed by reinforcement learning to improve performance. Experiments starting from Qwen3.5‑9B show that SCOPE‑RL achieves high task success and attack‑avoidance rates, outperforming other agents on OSWorld and OS‑BLIND benchmarks.
By Zeyu Kang, Zhenyun Yin, Yang Zhang, Shan He, Shanzhe Lei, Yanjiu Zhong, Xinquan Chen, Yuhong Wang
arXiv:2604. 15579v2 Announce Type: replace-cross Abstract: There is increasing interest in integrating AI agents that invoke tools into domain-specific commercial software, where unintended tool calls can cause serious security and safety incidents.
By Yining Hong, Yining She, Eunsuk Kang, Christopher S. Timperley, Christian K\"astner