arXiv AI By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter

Metacognitive Myopia in Large Language Models

Read the original on arXiv AI →

arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Beliefs and Behavior in Language Models

arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...

By Alex Smolin, Bryan Wilder
arXiv AI
Sep 2

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection

BiasGym is a cost‑effective, generalizable framework that injects specific biases into large language models via token‑based fine‑tuning while keeping the model frozen. It then uses two debiasing methods—Scope and Steer—to identify and suppress or redirect the components responsible for biased behavior. The framework enables consistent bias elicitation, precise localization of bias associations, and targeted debiasing without harming downstream performance, and it has been shown to reduce real‑world stereotypes such as labeling Italians as reckless drivers.

By Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein
arXiv AI
Sep 16

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

The paper investigates whether large language models (LLMs) possess intrinsic value systems and how to quantify and align them. By projecting responses from 106 LLMs and 95,000 human survey profiles into a shared sociological space, the authors confirm that LLMs do have values, though these values form a concentrated, idealized core rather than mirroring human diversity. They introduce the Prior-Environment-Cognition (PEC) framework to mathematically define value expression and propose an adaptive Alignment Prescription that identifies minimal interventions—ranging from prompts to targeted parameter updates—to steer LLM values efficiently without harming general performance.

By Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu, Nai Ding, Lai Jiang, Congyan Lang, Bing Li, Weiming Hu
arXiv AI
Sep 10

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g.,...

By Sunny Rai, Jinyi Kuang, Reyhan Jamalova, Annie Lou, Cristina Bicchieri, Niyati Malhotra, Victor Hugo Orozco-Olvera, Ana Maria Munoz-Boudet, Lyle H Ungar, Sharath C Guntuku
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.