Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
arXiv:2606. 12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.
arXiv:2607. 03932v1 Announce Type: cross Abstract: LLMs can be conveniently adapted to a diverse set of tasks, e.
arXiv:2606. 12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.
The paper investigates why prompt optimization works better for some tasks than others by decomposing reward variance into response variance and system‑prompt variance. It finds that optimization succeeds when system‑prompt variance dominates, and that adding more user prompts can actually reduce this variance, especially on heterogeneous datasets. To address this, the authors propose $p1$, a filtering method that selects a small set of high‑variance user prompts, which improves optimization on reasoning benchmarks and even allows a system prompt trained on just two AIME 24 prompts to generalize well.
arXiv:2606. 24841v1 Announce Type: new Abstract: Prompt-based learning has emerged as a dominant paradigm in natural language processing.
arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.
arXiv:2601. 02896v3 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.
HiVe is a prompt‑tuning framework that builds a hierarchy of prompts by exploiting inter‑task relationships during training. It uses a vertical mixture‑of‑experts (V‑MoE) at inference to compose prompts at the level of specialization needed for each input, allowing input‑dependent prompt adaptation. Experiments demonstrate that HiVe consistently outperforms strong prompt‑tuning baselines across diverse tasks.
arXiv:2512. 08724v3 Announce Type: replace Abstract: Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been repeatedly shown to exhibit social biases across sensitive attributes such as gender, race and age.
Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.
arXiv:2606.12234v2 Announce Type: replace Abstract: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the invol...
arXiv:2609.39927v1 Announce Type: new Abstract: Prompt optimization improves the performance of language-model systems on downstream tasks by refining their prompts. Classical methods evaluate prompt...
Hidden‑Shot introduces an implicit prompt mechanism that extracts task‑specific visual information and merges it with in‑task processing to boost one‑shot performance on new low‑level vision tasks. The method injects this prompt cost‑effectively while minimally altering the base generalist model’s architecture. A data‑driven evaluation framework, C/U assessment, is proposed to systematically test generalization across conventional and unconventional tasks, and experiments on seven and ten datasets show Hidden‑Shot outperforming state‑of‑the‑art models.
arXiv:2607. 14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.