Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By co...
arXiv:2602.13576v2 Announce Type: replace-cross
Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-langu...
By Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
JevOut demonstrates that natural, short additions to the context of decision models can flip their outputs from correct to incorrect, even when the correct answer remains unchanged. By optimizing context additions while keeping the source, question, choices, and gold answer fixed, the study found that 61.4% of initially correct decisions were redirected to a wrong option, with 45% receiving high confidence. Similar fragility was observed across three other decision systems on seven datasets, with flip rates between 64.9% and 73.2%.
By Zixiang Xu
arXiv:2607. 23976v1 Announce Type: cross Abstract: Appending a two-word confirmation tag to a decision question -- "Is X the better choice?
By Tapan Parikh
Appending a two-word confirmation tag to a decision question -- "Is X the better choice? " versus "X is the better choice, right?
arXiv:2609.39496v1 Announce Type: cross
Abstract: Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefine...
By Jike Zhong, Ming Li, Yuxiang Lai
The paper introduces Janus, a method for validating error patterns in language models by comparing error rates across predefined yes/no properties and using shuffled decoy labels to set significance thresholds. Janus requires that a pattern’s error difference surpasses the decoy-derived threshold and is replicated on held‑out data before reporting. Experiments on a controlled code‑finding task confirm several meaningful error patterns, while on MuSiQue and LongBench v2 Janus reports no confirmed patterns for the tested properties, contrasting with standard shuffling tests that sometimes confirm patterns.
By Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
arXiv:2609.16145v1 Announce Type: new
Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? W...
By Gautam Kishore
arXiv:2609.23886v1 Announce Type: new
Abstract: Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears wit...
By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv:2511. 00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks.
By Mina Taraghi, Yann Pequignot, Amin Nikanjam, Mohamed Amine Merzouk, Foutse Khomh