arXiv AI

The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models

arXiv:2606. 05183v2 Announce Type: replace-cross Abstract: Pass/fail safety evaluation reports whether a model refused.

arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil
arXiv AI
Aug 26

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The paper introduces a two‑dimensional construct validity framework for evaluating large language models (LLMs) as judges, defining invariance (S) and sensitivity (R) to construct‑preserving and construct‑changing edits. Experiments across seven judges and four domains reveal high invariance (average S = 0.945) but low sensitivity (average R = 0.319), with sensitivity varying by edit type. Audits of public label sets show that surface‑only predictors can reproduce a substantial portion of labels, underscoring that high agreement does not guarantee construct validity.

By Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
arXiv AI
Sep 2

Commit-first LLM judging inherits the judge's own errors

The paper investigates whether widely used evaluation frameworks for large language models (LLMs) implement a defense called commit‑first judging, which requires a judge to solve a task itself before accepting a candidate answer. Across 24 configurations in eight popular frameworks, none use the full commit‑first method; nine use a weaker variant that is ineffective. In controlled experiments, the weaker variant allowed systems to game the judge, while the full commit‑first approach eliminated this vulnerability but sometimes worsened evaluation when the judge’s own answer was incorrect.

By Idil Gozel
arXiv Computation and Language
3d ago

Safety Monitors Mostly Catch What the Model Already Refuses

The paper evaluates safety monitors by measuring recall only on prompts that the target model actually answers, rather than on all harmful prompts. Across several guard systems, recall at a 1% false‑positive rate drops sharply when focusing on answered prompts, with monitors catching refused requests 1.1–6.4 times more often than answered ones. Rewriting prompts to be less explicit dramatically increases compliance and reveals that many harmful requests slip past monitors, especially when phrasing is softened. Fine‑tuning guards on these rewritten prompts improves recall from 0.24 to 0.89 on answered requests and generalizes to unseen benchmarks.

By Sripad Karne
arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu