The paper presents a framework of formal operations for assembling context in large language model (LLM)-based engineering design, involving modular context units such as policy prompts, reference units with persistence, and user questions with prompt vectoring. It also introduces a formal method for evaluating modelling-as-code LLM outputs, assessing compliance to intent from LLM answers and the support LLMs provide for systems architecture modelling.
By Vinicius Kaster Marini, Petter Krus
arXiv:2607. 29389v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated a strong ability to generate syntactically correct code from natural-language specifications.
By Jan Marius St\"urmer, Jascha Knack, Tobias Koch, Andreas Weinmann
arXiv:2608. 05234v1 Announce Type: new Abstract: Building reliable applications that leverage large language models (LLMs) remains a significant challenge.
By Louis Mandel, Guillaume Baudart, Mandana Vaziri, Martin Hirzel
The paper introduces Program Executability Prediction (PrEx), a task that asks large language models (LLMs) to determine whether a program is semantically valid or invalid and, if invalid, to identify the violated formal rule. To evaluate this, the authors create a dataset of systematically generated invalid programs derived from valid ones and test open‑source coding LLMs across different semantic formalisms, semantic shifts, and program splits (human‑written, LLM‑translated, fuzzer‑generated). Results show that LLMs rely more on pre‑training priors than on the provided semantics, performing poorly on modified semantics and with increasing program complexity.
By Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric
The paper introduces FlowGen, a system that automates the construction of use case flows using large language models (LLMs). FlowGen extracts semantic elements via an LLM-based Semantic Information Processing module, builds a Semantic Relational Graph encoded by an enhanced R-GAT for basic flow generation (BFGen), and adds branch point prediction (BPP) and branch-conditioned alternative flow generation (AFGen). Experiments on 13 public and 7 industrial datasets show FlowGen outperforms baselines across precision, recall, F1, and AUC metrics for all three components.
By Guangyu Wang, Bangqi Li, Ji Wu, Zhijun Shao
arXiv:2607. 15845v2 Announce Type: replace Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions.
By Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu