arXiv:2606. 26185v1 Announce Type: new Abstract: LLM-as-judge ("grader") components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream deployment decisions.
By Hiroki Tamba
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
JevAdvBench introduces the first adversarial benchmark for reinforcement‑learning‑based calibrated decision (RLCD) models, providing 812 typed questions across 66 scenarios and a black‑box attack suite of 9,744 single‑edit variants. The benchmark evaluates attacks by comparing each perturbed decision to the model’s own clean decision and to an identical re‑run, revealing that rewording changes decisions by only 1.2 percentage points while certain injected opinions can flip 12.1% of decisions and lower confidence below 0.8 in 38% of cases. These findings demonstrate that RLCD models can be significantly misled by seemingly innocuous input edits, underscoring the need to treat the state as untrusted in applications.
By Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked.
arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.
By Yang Gao (Veyon Solutions)
arXiv:2609.38612v1 Announce Type: new
Abstract: As natural language drives more applications, language models increasingly run inside programs as decision components: the program sends them the curre...
By Jhen-Ke Lin, Chung Chun Wang
arXiv:2606. 09863v1 Announce Type: new Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise.
By Laksh Advani
AutoTuneBench introduces a trustworthy measurement protocol for evaluating how large language model agents auto‑tune GPU kernels and serving engines. The benchmark addresses four failure modes—strawman baselines, machine‑dependent timing, saturated tasks, and infrastructure defects—by enforcing code‑frozen protocols, database validation, anti‑cheat checks, pre‑registered comparisons, and external result anchoring. Using this protocol, the authors demonstrate that previously reported speedups are inflated, revealing more modest improvements across different engines and machines.
By Li Chen
The paper introduces Jev, a reinforcement‑learning‑trained model that provides calibrated probability answers to typed questions about a single input in one call. Jev is evaluated on RLCDAlignBench, a benchmark covering ten alignment failures across 44 tests and five target models, achieving a median AUROC of 0.886 zero‑shot and outperforming supervised baselines on most tasks. The study shows that question wording has little impact, while contextual fields that encode labels are more influential, and that Jev matches human‑label agreement while being 63× cheaper than LLM‑judge scorers.
By Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang
The paper evaluates safety monitors by measuring recall only on prompts that the target model actually answers, rather than on all harmful prompts. Across several guard systems, recall at a 1% false‑positive rate drops sharply when focusing on answered prompts, with monitors catching refused requests 1.1–6.4 times more often than answered ones. Rewriting prompts to be less explicit dramatically increases compliance and reveals that many harmful requests slip past monitors, especially when phrasing is softened. Fine‑tuning guards on these rewritten prompts improves recall from 0.24 to 0.89 on answered requests and generalizes to unseen benchmarks.
By Sripad Karne
arXiv:2609.22910v1 Announce Type: new
Abstract: Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the mod...
By Kunyu Peng, Junming Liu, Ruiqi He, Qingzhuo Wang, Jianzhong Qi, Xianhui Liu
The paper presents Baszta, a Polish multi‑label content‑safety classifier trained by fine‑tuning the 124M‑parameter allegro/herbert‑base‑cased model on five categories (hate, vulgarity, sexual content, crime, self‑harm) using a Focal + R‑Drop objective. In out‑of‑distribution evaluation on the Gadzi Język benchmark, Baszta achieves a small but statistically significant improvement in micro‑F1 over the Bielik Guard system, though the macro‑F1 advantage disappears when both models are properly tuned. The study also explores calibration techniques, showing that per‑category temperature scaling can recover performance lost by Platt scaling or isotonic regression, and discusses the trade‑offs between robust calibration and adversarial recall.
By Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski