arXiv AI By Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang

JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models

Read the original on arXiv AI →

JevAdvBench introduces the first adversarial benchmark for reinforcement‑learning‑based calibrated decision (RLCD) models, providing 812 typed questions across 66 scenarios and a black‑box attack suite of 9,744 single‑edit variants. The benchmark evaluates attacks by comparing each perturbed decision to the model’s own clean decision and to an identical re‑run, revealing that rewording changes decisions by only 1.2 percentage points while certain injected opinions can flip 12.1% of decisions and lower confidence below 0.8 in 38% of cases. These findings demonstrate that RLCD models can be significantly misled by seemingly innocuous input edits, underscoring the need to treat the state as untrusted in applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

The paper introduces Jev, a reinforcement‑learning‑trained model that provides calibrated probability answers to typed questions about a single input in one call. Jev is evaluated on RLCDAlignBench, a benchmark covering ten alignment failures across 44 tests and five target models, achieving a median AUROC of 0.886 zero‑shot and outperforming supervised baselines on most tasks. The study shows that question wording has little impact, while contextual fields that encode labels are more influential, and that Jev matches human‑label agreement while being 63× cheaper than LLM‑judge scorers.

By Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang
arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai