arXiv:2609.07943v1 Announce Type: new
Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...
By Alex Smolin, Bryan Wilder
BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.
By Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein
The paper investigates whether large language models (LLMs) possess intrinsic value systems and how to quantify and align them. By projecting responses from 106 LLMs and 95,000 human survey profiles into a shared sociological space, the authors confirm that LLMs do have values, though these values form a concentrated, idealized core rather than mirroring human diversity. They introduce the Prior-Environment-Cognition (PEC) framework to mathematically define value expression and propose an adaptive Alignment Prescription that identifies minimal interventions—ranging from prompts to targeted parameter updates—to steer LLM values efficiently without harming general performance.
By Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
By Anthony Baez, Sheer Karny, Pat Pataranutaporn
arXiv:2609.05437v1 Announce Type: new
Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g.,...
By Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.