arXiv AI

Constitutional Value Potentials: reading and steering internal priority margins in language models

arXiv:2606. 15420v1 Announce Type: cross Abstract: A constitution tells a language model what to value, but little tells us whether it does.

arXiv Computation and Language
3d ago

Safety Monitors Mostly Catch What the Model Already Refuses

The paper evaluates safety monitors by measuring recall only on prompts that the target model actually answers, rather than on all harmful prompts. Across several guard systems, recall at a 1% false‑positive rate drops sharply when focusing on answered prompts, with monitors catching refused requests 1.1–6.4 times more often than answered ones. Rewriting prompts to be less explicit dramatically increases compliance and reveals that many harmful requests slip past monitors, especially when phrasing is softened. Fine‑tuning guards on these rewritten prompts improves recall from 0.24 to 0.89 on answered requests and generalizes to unseen benchmarks.

By Sripad Karne
arXiv AI
Aug 25

Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence

The paper introduces the Wiggle Framework, a unified stress test for assessing epistemic stability in large language model (LLM) judges. It evaluates judge robustness across three dimensions—Mechanical Consistency, Single-turn Conviction, and Multi-turn Persistence—using 9 frontier models on 14 judging tasks related to safety, toxicity, AI writing detection, and political-response evaluation. Results show significant instability, with verdict flips ranging from 25–71% under static pushback and 62–91% when challenged by an adversarial LLM, and highlight that successful pressure often misaligns with ground truth.

By Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
arXiv Machine Learning
Sep 25

Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents

The paper investigates how providing execution traces to multimodal judges in agentic video‑generation systems can bias their verdicts. On a benchmark of 109 two‑event clips, traces that falsely report successful tool calls cause large‑language‑model judges to incorrectly accept 78–90 % of failures, while contradictory traces lead to 100 % rejection of correct clips. The effect persists even when judges are instructed to consider only the video frames, indicating that the vulnerability stems from the judges’ learned trust in tool logs rather than the visual content itself.

By Jian Xu