arXiv AI

Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations

arXiv:2604. 22207v2 Announce Type: replace-cross Abstract: Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processing.

arXiv AI
Sep 17

Which LLM is Best for Translating Natural Language Goals to PDDL

The paper evaluates how well current Large Language Models can translate natural language goals, written by video game testers, into well‑formed PDDL targets for classical planning. Using a carefully designed prompt template, six state‑of‑the‑art LLMs were tested on correctness, speed, and error tendencies with real‑world benchmarks. All models achieved high correctness (>92%), with Gemini 2.5 Flash reaching 96% accuracy and the fewest false positives, while GPT‑4.1 was the fastest, yet differences in performance and occasional failures due to ambiguity and domain limits remain.

By Tomas Balyo, Lukas Chrpa, G. Michael Youngblood
arXiv AI
Jun 17

PromptMN: Pseudo Prompting Language

arXiv:2606. 17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals, constraints, and expected outputs are often buried in prose or left implicit.

By Enkhzol Dovdon
arXiv AI
Aug 19

Parametric Knowledge in RAG-SFT for Domain-Specific Document Generation

The paper investigates Retrieval-Augmented Generation fine‑tuning (RAG‑SFT) for generating requirements documents in electronics engineering, comparing two 7B models trained with different data strategies. It introduces a claim‑based evaluation pipeline, C‑FEX, and a new metric, Parametric Knowledge Precision (PKP), to assess factuality of model‑generated claims. Results show that fine‑tuned 7B models can match or surpass a 72B baseline, but standard metrics may mislead, and fine‑tuning reduces hallucination by encouraging more reliable use of parametric knowledge.

By Julian Oestreich, Maximilian Bley, Frank Binder, Lydia M\"uller, Andr\'e Alcalde, Maksym Sydorenkoq
Hugging Face Trending Papers
Jul 29

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances.

arXiv Computation and Language
Sep 17

Relationally Guided Use Case Modeling with LLMs

The paper introduces FlowGen, a system that automates the construction of use case flows using large language models (LLMs). FlowGen extracts semantic elements via an LLM-based Semantic Information Processing module, builds a Semantic Relational Graph encoded by an enhanced R-GAT for basic flow generation (BFGen), and adds branch point prediction (BPP) and branch-conditioned alternative flow generation (AFGen). Experiments on 13 public and 7 industrial datasets show FlowGen outperforms baselines across precision, recall, F1, and AUC metrics for all three components.

By Guangyu Wang, Bangqi Li, Ji Wu, Zhijun Shao
arXiv Machine Learning
Aug 27

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

The paper introduces Multi-Objective In-context Knowledge Editing (MO‑IKE), a reinforcement learning framework that treats prompt construction for knowledge editing as a constrained Markov decision process. MO‑IKE jointly optimizes three competing objectives—reliability, generality, and specificity—by training a dynamic retriever to balance these goals and produce globally coherent prompts. Experiments on Llama‑3.2 show that MO‑IKE raises edit success from 85.0 % to 92.0 %, improves paraphrase consistency from 77 % to 79 %, and boosts retention rate by 23 % compared to earlier RL‑based methods.

By Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang, Haipeng Chen