arXiv AI By Zihan Chen, Di Zhu, Lei Zheng, Weiling Li

Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers

Read the original on arXiv AI →

The paper investigates how procedural traces—detailed step-by-step accounts of a language model’s reasoning—affect the decision-making of LLM overseers tasked with auditing another model’s outputs. Using signal detection theory, the authors evaluated five overseers on 19 compliance tasks, finding that while error detection remains high when disconfirming evidence is always visible, more elaborate traces shift the decision criterion toward rejection, leading to increased false alarms. The study also shows that providing option labels reduces the stated inability to link evidence to options, yet some overseers still exhibit residual rejection of correct work that grows with trace detail.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.

By Yuchen Han, Cheng Yan, Wuyang Zhang