arXiv AI

What Will Remain Human in Software Architecture? A Focus Group Report

The report examines how software architects view the growing use of AI development agents in their field. A focus group of 22 industry and academic participants discussed current practices, trust, validation, governance, and educational implications, concluding that decision‑making, accountability, and guardrail authoring remain human responsibilities. They introduced the concept of harness engineering—building systems that govern AI‑assisted creation—and identified criticality and cognitive debt as key factors for calibrating human oversight.

arXiv AI
Sep 1

Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act

The paper "Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act" outlines a detailed framework for implementing AI Regulatory Sandboxes (AIRS) under the EU AI Act. It maps the sandbox lifecycle into 29 activities, distinguishes between a Core AIRS and an Extended AIRS that includes an AI Technical Sandbox (AITS), and derives 15 infrastructural and governance requirements linked to these activities and provider obligations. The authors also introduce the Sandbox Configurator, an open‑source tool to instantiate AITS environments, aiming to provide structured workflows for regulators, robust evaluation methods for experts, and a transparent compliance pathway for AI providers.

By Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot
arXiv AI
Aug 26

AI Agents Push Humans Out of the Loop

AI agents are increasingly autonomous, posing significant risks that current designs hinder effective human oversight. The paper argues that oversight is degraded by both design choices and the cognitive decline of users who rely heavily on automation. It calls for prioritizing human cognitive needs in AI agent development, proposing design affordances and protocols to maintain critical judgment and counter skill atrophy.

By Margaret Mitchell, Avijit Ghosh, Samir Passi
arXiv Machine Learning
Sep 11

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

The paper argues that AI should be evaluated not only by principles but by concrete protocols that translate commitments into roles, requirements, records, oversight, and assessment. It introduces a rupture test linking institutional baselines to system evaluation, and distinguishes evidence‑bounded deployment from measurement‑bounded governance. The authors propose the RISE AI architecture to make bounded, evidence‑based claims about Responsibility, Inclusivity, Safety, and Empowerment, emphasizing the need for engineering, institutional repair, and ongoing moral judgment.

By Nitesh V. Chawla, Paulo Benanti
arXiv AI
Sep 15

AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems

The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.

By Murat Kantarcioglu
arXiv AI
4d ago

Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development

The paper examines how a large embedded systems company is transitioning to an AI‑first organization, focusing on the role of autonomous AI agents in software engineering. Through a mixed‑method study involving 40 workshop participants—scrum masters, architects, managers, and product owners—the authors identify expected impacts on team structure, required competencies, organizational strategies, and developer roles. The study concludes with a concrete roadmap and discusses implications for federated AI team formation, human‑in‑the‑loop practices, and sustainable AI adoption in embedded software engineering.

By Viktor Kjellberg, Srijita Basu, Simin Sun, Farnaz Fotrousi, Miroslaw Staron