arXiv:2608. 16627v1 Announce Type: cross Abstract: Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL).
By Mahdi Dhaini, Adam Dejl, Juraj Vladika, Volkan \"Ozer, Barbara Plank, Gjergji Kasneci
The paper investigates how small lexical changes in prompts can cause large performance swings in large language models. Using a dataset of 132,000 prompt variants, the authors uncover a scaling law linking higher average task performance to lower variance and greater robustness. They identify domain-specific terminology and explicit action directives as key linguistic factors that stabilize prompts, and propose an automated Prompt-Refining Agent that reduces performance variance by 40.7% in code generation while maintaining or improving mean performance.
By Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
arXiv:2602.21223v2 Announce Type: replace
Abstract: It is not only what we ask large language models (LLMs) to do that matters, but also how we ask them. Phrases like ``This is urgent'' or ``As your...
By Yilin Geng, Omri Abend, Eduard Hovy, Lea Frermann
arXiv:2606. 04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions.
By Akanksha Narula, Mofasshara Binte Rafique, Laurent Bindschaedler
Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.
By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
Instruction tuning is often thought to give language models a universal ability to follow instructions, but this study shows otherwise. By probing nine tasks across three models, the authors find that general probes reveal selective, not uniform, deficits, cross‑task transfer is weak and skill‑similar, and causal ablation uncovers sparse, asymmetric dependencies. The results suggest instruction following is a coordinated use of diverse linguistic skills rather than a single shared mechanism.
By Elisabetta Rocchetti, Alfio Ferrara