When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 28863v1 Announce Type: cross Abstract: AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts.
The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.
arXiv:2608. 04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments.
arXiv:2608. 00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.
arXiv:2605. 16281v2 Announce Type: replace-cross Abstract: Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident reporting, and impact assessment obligations are visible when AI systems fail.
arXiv:2607. 29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation.