The paper introduces Jev, a reinforcement‑learning‑trained model that provides calibrated probability answers to typed questions about a single input in one call. Jev is evaluated on RLCDAlignBench, a benchmark covering ten alignment failures across 44 tests and five target models, achieving a median AUROC of 0.886 zero‑shot and outperforming supervised baselines on most tasks. The study shows that question wording has little impact, while contextual fields that encode labels are more influential, and that Jev matches human‑label agreement while being 63× cheaper than LLM‑judge scorers.
By Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
By Chenyu Zhou
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
arXiv:2607. 18966v1 Announce Type: new Abstract: Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective.
By Axel H{\o}jmark, J\'er\'emy Scheurer, Evgenia Nitishinskaya, Felix Hofst\"atter, Jason Wolfe, Theodore Ehrenborg, Bronson Schoen, Alexander Meinke
The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.
By Zhuoang Cai