arXiv AI

Code Owns the Simulation, Jev Owns the Evaluation

arXiv:2610. 01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option.

arXiv AI
Sep 28

JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models

JevAdvBench introduces the first adversarial benchmark for reinforcement‑learning‑based calibrated decision (RLCD) models, providing 812 typed questions across 66 scenarios and a black‑box attack suite of 9,744 single‑edit variants. The benchmark evaluates attacks by comparing each perturbed decision to the model’s own clean decision and to an identical re‑run, revealing that rewording changes decisions by only 1.2 percentage points while certain injected opinions can flip 12.1% of decisions and lower confidence below 0.8 in 38% of cases. These findings demonstrate that RLCD models can be significantly misled by seemingly innocuous input edits, underscoring the need to treat the state as untrusted in applications.

By Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang
arXiv AI
Sep 25

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.

By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen