arXiv AI

Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design

The paper presents a framework of formal operations for assembling context in large language model (LLM)-based engineering design, involving modular context units such as policy prompts, reference units with persistence, and user questions with prompt vectoring. It also introduces a formal method for evaluating modelling-as-code LLM outputs, assessing compliance to intent from LLM answers and the support LLMs provide for systems architecture modelling.

arXiv AI
Sep 24

Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems

The paper introduces Constraint-Driven Context Engineering (CDCE), a design approach that treats domain constraints as primary drivers for creating AI system interfaces. CDCE identifies, characterises, and operationalises constraints to determine necessary context assets and their representations, improving the quality and domain appropriateness of AI-generated solutions. A comparative multiple‑case study across education, healthcare, and finance demonstrates CDCE’s applicability and shows how constraint characteristics shape the resulting interfaces.

By Xiwei Xu, Chen Wang, Mengmeng Yang, Yipeng Zhang, Jacky Jiang, Suyu Ma, Youyang Qu, Ming Ding, Liming Zhu
arXiv AI
Aug 18

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

arXiv:2608. 15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve.

By Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
arXiv AI
Aug 5

Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations

arXiv:2604. 22207v2 Announce Type: replace-cross Abstract: Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processing.

By Anna Arnaudo, Riccardo Coppola, Maurizio Morisio, Flavio Giobergia, Andrea Bioddo, Angelo Bongiorno, Luca Dadone
Hugging Face Trending Papers
Jul 7

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs).

Hugging Face Trending Papers
Jul 28

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: This study aims to eliminate logical inconsistencies and enforce structural conformance in LLM-generated requirements while quantifying the LLM's pre-validation decision uncertainty within a formal domain model.

arXiv Computation and Language
Sep 17

Relationally Guided Use Case Modeling with LLMs

The paper introduces FlowGen, a system that automates the construction of use case flows using large language models (LLMs). FlowGen extracts semantic elements via an LLM-based Semantic Information Processing module, builds a Semantic Relational Graph encoded by an enhanced R-GAT for basic flow generation (BFGen), and adds branch point prediction (BPP) and branch-conditioned alternative flow generation (AFGen). Experiments on 13 public and 7 industrial datasets show FlowGen outperforms baselines across precision, recall, F1, and AUC metrics for all three components.

By Guangyu Wang, Bangqi Li, Ji Wu, Zhijun Shao
arXiv AI
Jun 17

PromptMN: Pseudo Prompting Language

arXiv:2606. 17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals, constraints, and expected outputs are often buried in prose or left implicit.

By Enkhzol Dovdon