The paper "Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act" outlines a detailed framework for implementing AI Regulatory Sandboxes (AIRS) under the EU AI Act. It maps the sandbox lifecycle into 29 activities, distinguishes between a Core AIRS and an Extended AIRS that includes an AI Technical Sandbox (AITS), and derives 15 infrastructural and governance requirements linked to these activities and provider obligations. The authors also introduce the Sandbox Configurator, an open‑source tool to instantiate AITS environments, aiming to provide structured workflows for regulators, robust evaluation methods for experts, and a transparent compliance pathway for AI providers.
By Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot
The paper introduces a unified evaluation framework for assessing the trustworthiness of large language models, agentic AI, and multimodal systems. It connects output-level, trajectory-level, and cross-modal assessments across eight dimensions—capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency—while preserving system-specific metrics and providing uncertainty estimates. A meta-evaluation layer checks the validity, reliability, and reproducibility of the evaluation itself, and the framework aligns with governance standards and regulatory requirements.
By Shaina Raza, Ahmed Y. Radwan, Imran Liaquat, Kathryn Hume
arXiv:2607. 16130v1 Announce Type: cross Abstract: AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way.
By Andrea Ferrario
arXiv:2607. 14673v1 Announce Type: new Abstract: Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale.
By Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee
arXiv:2606. 01417v1 Announce Type: new Abstract: Turkey's e-Government Gateway (e-Devlet) serves over 68 million registered users with more than 9,200 government services, and is increasingly integrating artificial intelligence into citizen-facing applications such as chatbot assistants and eligibility assessments.
By Ahmet Kaplan
The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.
By Samira Maghool, Paolo Ceravolo
arXiv:2608. 04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments.
By Leah Davis, Dominic Martin, AJung Moon
arXiv:2609.37109v1 Announce Type: cross
Abstract: The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Establ...
By Idoia Landa-Oregi, Tom Deckenbrunnen, Alessio Buscemi, Daniele Pagani, German Castignani
The study examines how practitioners in AI-driven systems define, assess, and manage data quality, revealing six key themes. It highlights shifts in traceability, the use of models as quality assessors, and the emergence of new data objects such as agent context and synthetic data. The research proposes a lifecycle assurance framework to provide evidence that data supports specific AI claims throughout model behavior, judgments, and agent actions.
By Hariharan Gopinath, Jan Bosch, Helena Holmstr\"om Olsson
arXiv:2606. 00047v1 Announce Type: cross Abstract: Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used during training.
By Arthur Goemans, Dan Altman, Noemi Dreksler, Jonas Freund, Milan Gandhi, Zhengdong Wang, Sarah Cogan, Sebastien Krier, Demetra Brady, Lewis Ho, Allan Dafoe
arXiv:2607. 15480v1 Announce Type: new Abstract: As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority.
By Michael Papademas, Xenia Ziouvelou, Kostas Karpouzis, Vangelis Karkaletsis
arXiv:2608. 02679v1 Announce Type: cross Abstract: Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, tenant separation, and transparent workflows.
By Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo, Md Mahade Hasan, Kai-Kristian Kemell, Jussi Rasku, Mika Saari, Vilma Saari, Roope Pajasmaa, Markku Oivo, Pekka Abrahamsson