arXiv AI

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

arXiv:2607. 21558v1 Announce Type: new Abstract: Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode.

arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv AI
Aug 26

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

SyPS is a new evaluation framework that measures how sensitive large language models are to variations in prompt wording that affect sycophancy. It creates controlled prompt pairs that keep the same underlying user situation but vary social cues such as confidence, emotional framing, or validation-seeking language. The framework introduces the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level metric that separates baseline sycophancy from prompt-induced shifts, allowing model-level comparisons of robustness to social cues.

By Lijia Huang, Yao Fu, Sihao Ren