arXiv:2608. 06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong.
By Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
arXiv:2606. 24267v1 Announce Type: cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing.
By Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques
arXiv:2606. 24267v2 Announce Type: replace-cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing.
By Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques
Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By co...
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting.
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy.
arXiv:2606. 29718v1 Announce Type: cross Abstract: Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon tasks.
By Shijie Xia, Yikun Wang, Zhen Huang, Pengfei Liu
The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.
By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang
arXiv:2609.34284v2 Announce Type: replace
Abstract: Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its appl...
By Haeun Jang, Yonghyun Jun, Hwanhee Lee
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon tasks. The concern that increasing context length degrades model capabilities, known as context rot, has become a central issue for these applications.
arXiv:2609.39929v1 Announce Type: cross
Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and disting...
By Yuyang Wu, Yufeng Du, Hao Peng