Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel.
The paper introduces a claim‑safe protocol for evaluating closed‑loop AI systems, consisting of three actions: Refuse, Decompose, and Refresh. It demonstrates the protocol in a simulator with 24 policy components and 1,440 held‑out cases, showing that abstention and stable false admission rates are low while providing detailed statistical diagnostics. The approach emphasizes that evaluation results should be tied to observable support and statistical calibration rather than a single PASS/FAIL label.
By Peiying Zhu, Sidi Chang
The paper reports a preregistered audit of language‑model judges used as measurement instruments, revealing that the assumption that a model’s responses remain stable over time is invalid. Across nearly 53,000 audited requests, repeat rankings and byte‑identical replays fell far below required reliability thresholds, with three identified mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—explaining the discrepancy. The study proposes a three‑level snapshot‑identity framework, eight design rules, and a reporting checklist to prevent such reliability failures in future evaluations.
By Haoyaun Zhu, Jie Zhang
arXiv:2609.22582v1 Announce Type: new
Abstract: End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a rank reports an outcome, not the behaviour behind...
By Ruolin Yang, Zilin Huang, Buoyue Wang, Zhengyang Wan, Yuhao Luo, Zihao Sheng, Sikai Chen
The paper investigates the reliability of language‑model judges used as measurement instruments on shared endpoints. Through two preregistered audits of 52,988 requests, the authors found that repeat rankings and byte‑identical replays fell far short of required thresholds, revealing significant instability. They identify three mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—that explain the gap, and propose a snapshot‑identity ladder, design rules, and a reporting checklist to mitigate such failures.
The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.
By Elias Malomgr\'e, Pieter Simoens