arXiv AI

Benchmarking candidate coverage and rejection policy transfer in typed decision models

arXiv AI
4d ago

Benchmarking Candidate Coverage in Typed Decision Models

The paper introduces a paired candidate‑coverage benchmark protocol for typed decision models, evaluating two models—Laya and Jev—on datasets such as AG News, DBpedia, Emotion, and TREC. It reports that Laya detects a high percentage of missing-answer cases but also falsely rejects many valid candidates, whereas Jev shows lower false rejection rates but also lower detection of missing answers. The study highlights the need for separate measurements of classification, score ranking, and rejection policies, noting that the benchmark is descriptive and limited to reference‑label omission.

By Jiawen Lu, Tongtong Wu
arXiv AI
Sep 25

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

The paper introduces Jev, a reinforcement‑learning‑trained model that provides calibrated probability answers to typed questions about a single input in one call. Jev is evaluated on RLCDAlignBench, a benchmark covering ten alignment failures across 44 tests and five target models, achieving a median AUROC of 0.886 zero‑shot and outperforming supervised baselines on most tasks. The study shows that question wording has little impact, while contextual fields that encode labels are more influential, and that Jev matches human‑label agreement while being 63× cheaper than LLM‑judge scorers.

By Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang
arXiv AI
Sep 18

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.

By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie