arXiv:2511. 19829v3 Announce Type: replace Abstract: Prompt optimization has become a central mechanism for eliciting strong performance from LLMs, and recent work has made substantial progress by proposing diverse prompt evaluation metrics and optimization strategies.
By Ke Chen, Yifeng Wang, Hassan Almosapeeh, Haohan Wang
The paper investigates whether a frozen large language model can be personalized to individual users via prompt-space meta‑learning. Using the Muse framework, the authors evolve a shared adaptation prompt across a meta‑train user population and test it zero‑shot on over 200 held‑out users in two personalization benchmarks (LaMP‑2 and LaMP‑3). The results show that Muse does not outperform its un‑evolved seed prompt or a control that trains on mismatched user‑support pairs, and it is outperformed by simple few‑shot retrieval on the rating task. The authors attribute this failure to a meta‑objective collapse, where the validation objective is invariant to genuine user‑support correspondence, leading to over‑optimization of instruction polish rather than transferable adaptation.
By Liam Byrne, David Dylan, Orla Fitzgerald, Eoin Doyle, Ciara Nolan, Padraig Lynch, Sinead Gallagher
arXiv:2609.39882v1 Announce Type: new
Abstract: Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-...
By Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou
arXiv:2601. 02896v3 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.
By Harshvardhan Saini, Yiming Tang, Dianbo Liu
The paper investigates how small lexical changes in prompts can cause large performance swings in large language models. Using a dataset of 132,000 prompt variants, the authors uncover a scaling law linking higher average task performance to lower variance and greater robustness. They identify domain-specific terminology and explicit action directives as key linguistic factors that stabilize prompts, and propose an automated Prompt-Refining Agent that reduces performance variance by 40.7% in code generation while maintaining or improving mean performance.
By Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
The paper investigates why prompt optimization works better for some tasks than others by decomposing reward variance into response variance and system‑prompt variance. It finds that optimization succeeds when system‑prompt variance dominates, and that adding more user prompts can actually reduce this variance, especially on heterogeneous datasets. To address this, the authors propose $p1$, a filtering method that selects a small set of high‑variance user prompts, which improves optimization on reasoning benchmarks and even allows a system prompt trained on just two AIME 24 prompts to generalize well.
By Zhaolin Gao (Sid), Yu (Sid), Wang, Bo Liu, Thorsten Joachims, Kiant\'e Brantley, Wen Sun
arXiv:2608. 10471v1 Announce Type: new Abstract: Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals.
By Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala
arXiv:2511. 19829v2 Announce Type: replace Abstract: Most prompt-optimization methods refine a single static template, making them ineffective in complex and dynamic user scenarios.
By Ke Chen, Yifeng Wang, Hassan Almosapeeh, Haohan Wang
arXiv:2608. 01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts.
By B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S
Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.
By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
The paper introduces PersonaLink, a training‑free method that distills a user’s interaction history into a bounded three‑field persona and iteratively refines it by self‑evaluating a frozen 7B language model on held‑out labeled data. Each refinement rewrites the persona only if it does not regress on that slice, ensuring the persona remains bounded and query‑independent. On a 200‑user news categorization task (LaMP‑2), PersonaLink achieves 0.745–0.755 accuracy, statistically indistinguishable from BM25 retrieval’s 0.760–0.765 accuracy, demonstrating that distilled personas can match retrieval for classification but not for regression tasks.
By JaeHa Yoon, Minjun Park, Seoyeon Kim, Jiwoo Lee, Hyunwoo Choi, Dohyun Kang
arXiv:2606. 27226v1 Announce Type: new Abstract: Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug.
By Sangwoo Cho, Kushal Chawla, Pengshan Cai, Zefang Liu, Chenyang Zhu, Shi-Xiong Zhang, Sambit Sahu