arXiv AI
Sep 25

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.

By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang
arXiv Computation and Language
Sep 25

JevOut: Natural Context Can Flip Decision Models

JevOut demonstrates that natural, short additions to the context of decision models can flip their outputs from correct to incorrect, even when the correct answer remains unchanged. By optimizing context additions while keeping the source, question, choices, and gold answer fixed, the study found that 61.4% of initially correct decisions were redirected to a wrong option, with 45% receiving high confidence. Similar fragility was observed across three other decision systems on seven datasets, with flip rates between 64.9% and 73.2%.

By Zixiang Xu