The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.
By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv:2603. 13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring.
By Petter T\"ornberg
arXiv:2505. 24539v4 Announce Type: replace-cross Abstract: We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs).
By Celia Cintas, Miriam Rateike, Erik Miehling, Elizabeth Daly, Skyler Speakman
The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.
By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
By Anthony Baez, Sheer Karny, Pat Pataranutaporn
arXiv:2609.15998v1 Announce Type: new
Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property...
By Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds