Graphical-Probabilistic Modeling of Generative Flows in LLM-Native Software Systems
arXiv:2606. 15943v1 Announce Type: cross Abstract: Engineering LLM-native software remains a challenging and immature field.
The paper presents a framework of formal operations for assembling context in large language model (LLM)-based engineering design, involving modular context units such as policy prompts, reference units with persistence, and user questions with prompt vectoring. It also introduces a formal method for evaluating modelling-as-code LLM outputs, assessing compliance to intent from LLM answers and the support LLMs provide for systems architecture modelling.
arXiv:2606. 15943v1 Announce Type: cross Abstract: Engineering LLM-native software remains a challenging and immature field.
The paper introduces Constraint-Driven Context Engineering (CDCE), a design approach that treats domain constraints as primary drivers for creating AI system interfaces. CDCE identifies, characterises, and operationalises constraints to determine necessary context assets and their representations, improving the quality and domain appropriateness of AI-generated solutions. A comparative multiple‑case study across education, healthcare, and finance demonstrates CDCE’s applicability and shows how constraint characteristics shape the resulting interfaces.
arXiv:2607. 29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications.
arXiv:2607. 05985v1 Announce Type: new Abstract: This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation.
arXiv:2608. 15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve.
arXiv:2604. 22207v2 Announce Type: replace-cross Abstract: Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processing.
This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs).
arXiv:2609.36788v1 Announce Type: new Abstract: Incorporating rich task-relevant context, such as domain knowledge and external observations, is a key capability yet remains challenging for Bayesian...
Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: This study aims to eliminate logical inconsistencies and enforce structural conformance in LLM-generated requirements while quantifying the LLM's pre-validation decision uncertainty within a formal domain model.
arXiv:2607. 26220v1 Announce Type: cross Abstract: Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees.
The paper introduces FlowGen, a system that automates the construction of use case flows using large language models (LLMs). FlowGen extracts semantic elements via an LLM-based Semantic Information Processing module, builds a Semantic Relational Graph encoded by an enhanced R-GAT for basic flow generation (BFGen), and adds branch point prediction (BPP) and branch-conditioned alternative flow generation (AFGen). Experiments on 13 public and 7 industrial datasets show FlowGen outperforms baselines across precision, recall, F1, and AUC metrics for all three components.
arXiv:2606. 17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals, constraints, and expected outputs are often buried in prose or left implicit.