arXiv:2606. 14923v1 Announce Type: new Abstract: As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates.
By Yujiao Chen
arXiv:2606. 28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming.
By Charles L. Wang, Keir Dorchen, Peter Jin
arXiv:2605. 28114v2 Announce Type: replace Abstract: Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software.
By Messi H. J. Lee
arXiv:2608. 07070v1 Announce Type: cross Abstract: With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, undermining information credibility and social trust.
By Qin Li, Gui Zhang, Minyu Feng, Matjaz Perc, Attila Szolnoki
arXiv:2607. 15992v1 Announce Type: new Abstract: Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings.
By Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, Mar\'ia Llorente S\'anchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman
The paper "Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement" highlights a gap in enterprise AI governance, where 78% of organizations lack auditable evidence of policy enforcement. It introduces AGIL, a five-layer adaptive governance architecture that uses machine learning for real-time detection, risk classification, sub-100ms policy enforcement, continuous attestation, and policy evolution. The authors argue that the failure is organizational and architectural, not technical, and call for future empirical validation of AGIL.
By Sandeep Bokkasam, B. Durgalakshmi
The paper introduces the concept of Evolutionary Safety for recursive self-improving AI, focusing on how safety properties evolve as an AI system and its successors change. It identifies key risks such as intent drift, error accumulation, and safety-property erosion, and presents a taxonomy covering agent state, model state, evaluation, environment, and update mechanisms. The authors propose methods for discovering and evaluating evolutionary risks, and outline governance principles for modification, selection, authorization, provenance, and recovery, while highlighting open problems for maintaining safety in persistent, adaptive, and recursively self-improving systems.
By Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
arXiv:2607. 19292v1 Announce Type: cross Abstract: Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios.
By Gjergji Kasneci, Enkelejda Kasneci
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development.
Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness.
The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.
By Samira Maghool, Paolo Ceravolo
arXiv:2605. 02640v2 Announce Type: replace Abstract: As artificial intelligence (AI), including machine learning (ML) models and foundation models (FMs), are increasingly deployed in high-stakes domains, ensuring their trustworthiness has become a central challenge.
By Ruta Binkyte, Ivaxi Sheth, Zhijing Jin, Mohammad Havaei, Bernhard Sch\"olkopf, Mario Fritz