Semantic Feature Analysis (SFA) is a method that refines agent specifications without performing any rollout-based search. It analyzes existing execution traces, clusters workflow node outputs, extracts semantic feature classes via an extended subject‑verb‑object schema, ranks these features with a decision tree, and injects the most impactful features back into the system prompt. Evaluations on four benchmarks show that SFA consistently outperforms five prompt‑optimisation algorithms and a single‑reflection baseline, especially when rollout costs are high.
By Yuval David, Fabiana Fournier, Lior Limonad, Hadar Mulian
arXiv:2609.39927v1 Announce Type: new
Abstract: Prompt optimization improves the performance of language-model systems on downstream tasks by refining their prompts. Classical methods evaluate prompt...
By Junyang Chen, Zecheng Wang, Jingbang Chen
arXiv:2607. 14105v1 Announce Type: cross Abstract: For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints.
By Cedric Richter, Salah Ghamizi, Mike Papadakis
Naive Prompt Optimization (NPO) is a lightweight, single‑lineage method that iteratively refines prompts using a teacher model’s rollout feedback. It matches or surpasses the performance of more complex optimizers like GEPA while requiring fewer rollouts, and its advantage grows with stronger teacher models. In interactive games, NPO remains competitive, and prompts optimized by NPO transfer well to other student models within the same family.
By Yuan Chang, Xiaoqi Chen
ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by separating optimization into Diagnose, Propose, and Select phases. It clusters training errors, generates diverse candidates, and applies bootstrap stability selection, achieving a 3.76‑point accuracy gain over GEPA on seven NLP benchmarks while producing 47% shorter prompts. Cross‑model tests on four additional student models confirm ESPO’s superior average accuracy, notably improving Qwen3 GSM8K from 15.00% to 91.40%.
arXiv:2606. 19605v1 Announce Type: cross Abstract: Multi-step LLM pipelines fail through interactions among retrieval, reasoning, and formatting steps, so prompt-only optimization can miss bottlenecks in the chain.
By Paul Kassianik, Baturay Saglam, Huaibo Zhao, Blaine Nelson, Supriti Vijay, Aman Priyanshu, Amin Karbasi
arXiv:2609.40361v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based o...
By Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang
ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by splitting the optimization process into Diagnose, Propose, and Select phases. It clusters training errors into structural patterns, generates diverse candidate prompts through four complementary strategies, and applies bootstrap stability selection. Across seven NLP benchmarks, ESPO improves average accuracy by +3.76 pp over GEPA, produces prompts 47 % shorter, and achieves higher accuracy on four additional student models, with the largest gain on Qwen3 GSM8K.
By Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.
By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
arXiv:2607. 25675v1 Announce Type: new Abstract: Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box.
By Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao
arXiv:2608.16831v2 Announce Type: replace
Abstract: Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current p...
By Minh-Ha Nguyen, Ngoc-Ngo Quang Tran, Thuy Dung Nguyen, Cathy Shyr
arXiv:2605. 29668v2 Announce Type: replace Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment.
By Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert, Lisa Adams, Keno Bressem