arXiv AI By Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques

Pigeonholing: Bad prompts hurt models to collapse and make mistakes

Read the original on arXiv AI →

arXiv:2606. 24267v1 Announce Type: cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 14

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy.

arXiv AI
3d ago

Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits

arXiv:2609.35356v2 Announce Type: replace Abstract: Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted gen...

By Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan Panigrahi, Srishti Gureja, Helen Yannakoudakis, Robert Mullins, Victor Gillioz, Daniel Tan, Maxime Rich\'e
arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai