arXiv Computation and Language

JevOut: Natural Context Can Flip Decision Models

JevOut demonstrates that natural, short additions to the context of decision models can flip their outputs from correct to incorrect, even when the correct answer remains unchanged. By optimizing context additions while keeping the source, question, choices, and gold answer fixed, the study found that 61.4% of initially correct decisions were redirected to a wrong option, with 45% receiving high confidence. Similar fragility was observed across three other decision systems on seven datasets, with flip rates between 64.9% and 73.2%.

Hugging Face Trending Papers
Jul 14

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy.

arXiv AI
Sep 25

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.

By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang