Prompt engineering helps you write better prompts—but it doesn’t help you change them safely. This article explores a common production failure where a simple variable rename breaks every live call, and introduces a lightweight static analysis tool that treats prompts like contracts, catching breaking changes before they ship.
By Emmimal P Alexander
But don't let the model check itself The post Design Loops, Not Prompts appeared first on Towards Data Science .
By Javier Marin
arXiv:2607. 28617v2 Announce Type: replace Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications.
By Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
arXiv:2607. 06074v1 Announce Type: cross Abstract: Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature.
By Rohit Mehra, Kapil Singi, Vikrant Kaulgud, Vibhu Saujanya Sharma, Swapnajeet Gon Choudhury, Swati Sharma, Adam P. Burden, Majd Sakr
The article describes how the author constructed a prompt dependency graph to identify which prompts are affected when a single prompt changes. By separating all reachable components from the smaller subset that truly requires evaluation, the graph helps focus retesting efforts. This approach streamlines testing by pinpointing only the prompts that need targeted evaluation.
By Emmimal P Alexander
The paper investigates how combining soft prompts via task arithmetic can reduce reliance on confounding variables in classification models. It introduces Hybrid Prompt Arithmetic (HyPA), which merges task prompts with linearized confounder prompts to counteract spurious correlations. Experiments across multiple benchmarks show that HyPA consistently improves the robustness‑performance trade‑off under distribution shift, and analysis of hidden representations suggests it mitigates confounding by diminishing the influence of confounder signals.
By Zhecheng Sheng, Yongsen Tan, Xiruo Ding, Trevor Cohen, Serguei Pakhomov
arXiv:2609.15045v1 Announce Type: cross
Abstract: Prompt echoing is a recognized failure mode of instruct language models, in which a model instead of generating a response, mirrors the provided prom...
By Inez Okulska, Bartosz Naskr\k{e}cki, Jan Piotrowski, Tomasz Steifer
arXiv:2608. 11513v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code.
By Alex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodr\'iguez-P\'erez
A practical tutorial for recording model tool requests, real function results, patches, checks, screenshots, and a saved run log. The post How to Debug AI Coding Agents When They Change the Wrong Thing appeared first on Towards Data Science .
By Abdullahi Dattijo
Enterprise Document Intelligence [Vol. 1 #9bis] - Your RAG isn’t hallucinating, it’s answering the wrong context faithfully.
By angela shi
arXiv:2506.13182v3 Announce Type: replace-cross
Abstract: [...] Since then, various APR approaches, especially those leveraging the power of large language models (LLMs), have been rapidly developed...
By Anh Ho, Thanh Le-Cong, Bach Le, Christine Rizkallah
arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.
By Valentin No\"el