arXiv AI

Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems

The paper introduces Constraint-Driven Context Engineering (CDCE), a design approach that treats domain constraints as primary drivers for creating AI system interfaces. CDCE identifies, characterises, and operationalises constraints to determine necessary context assets and their representations, improving the quality and domain appropriateness of AI-generated solutions. A comparative multiple‑case study across education, healthcare, and finance demonstrates CDCE’s applicability and shows how constraint characteristics shape the resulting interfaces.

arXiv AI
Sep 11

Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design

The paper presents a framework of formal operations for assembling context in large language model (LLM)-based engineering design, involving modular context units such as policy prompts, reference units with persistence, and user questions with prompt vectoring. It also introduces a formal method for evaluating modelling-as-code LLM outputs, assessing compliance to intent from LLM answers and the support LLMs provide for systems architecture modelling.

By Vinicius Kaster Marini, Petter Krus
arXiv AI
Sep 25

The Gold in Bias: Maturing the AI Design Process through Verification

The paper proposes rethinking bias in AI as a diagnostic tool rather than merely a flaw to be minimized. It introduces a multidimensional framework that examines bias across origin, lifecycle emergence, technical causes, and validation methods, covering 30 bias types, 16 verification methods, and 20 countermeasures for both traditional and generative AI. The authors present a hierarchical evidence framework distinguishing internal and external validity, and advocate for Ethics by Design principles to embed bias verification throughout the AI development lifecycle.

By Samira Maghool, Paolo Ceravolo
arXiv AI
Jun 15

Thinking Outside the [Chat]Box: Bridging Computer Science and Industrial Design for Cognitive-Inclusive Generative AI

arXiv:2606. 14306v1 Announce Type: cross Abstract: Current Generative AI (GenAI) interfaces remain largely constrained to chatbox interaction, which can impose high cognitive demands on users and create substantial barriers for people with intellectual disabilities (ID), including prompt formulation difficulties, response overload, and limited mechanisms to assess information reliability.

By Virginia Francisco, Daniel Guasch, Raquel Herv\'as
arXiv AI
Sep 15

AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems

The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.

By Murat Kantarcioglu
arXiv Computation and Language
Aug 28

Agent Seer: Synthesizing Scenarios from Specification Understanding

Agent Seer is a pipeline that automatically synthesizes realistic evaluation scenarios for AI agents that use external tools, using only the tool’s specification (function names, natural‑language descriptions, and typed parameter schemas). Starting from a single Model Context Protocol (MCP) specification, it enriches raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock‑data‑grounded multi‑turn dialogues that demonstrate strong tool‑calling correctness and conversational coherence. Across seven diverse MCP specifications, the pipeline achieves high quality, with parameter‑schema complexity emerging as the main driver of quality variation and argument‑value accuracy identified as the dominant failure mode.

By Harish Karumuri, Mahesh Vemula, David Lopes Pegna
Hugging Face Trending Papers
Jul 28

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: This study aims to eliminate logical inconsistencies and enforce structural conformance in LLM-generated requirements while quantifying the LLM's pre-validation decision uncertainty within a formal domain model.

arXiv AI
Jun 17

PromptMN: Pseudo Prompting Language

arXiv:2606. 17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals, constraints, and expected outputs are often buried in prose or left implicit.

By Enkhzol Dovdon
arXiv Machine Learning
Jul 1

From Failure to Alignment: A Requirements Engineering Framework for Machine Learning Systems

arXiv:2606. 31589v1 Announce Type: cross Abstract: Organisations designing, developing, and deploying machine learning systems (MLS) need to be able to check that these systems are trustworthy, and communicate this clearly to their stakeholders, be they different categories of users, engineers, or wider society.

By Amel Bennaceur, Gopi Krishnan Rajbahadur, Prince Mercy, Bashar Nuseibeh, Faeq Alrimawi