arXiv Machine Learning

STEREODISCO: Discovering Stereotypicality in LLMs

arXiv:2607. 27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes.

arXiv AI
Jul 29

Localizing Persona Representations in LLMs

arXiv:2505. 24539v4 Announce Type: replace-cross Abstract: We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs).

By Celia Cintas, Miriam Rateike, Erik Miehling, Elizabeth Daly, Skyler Speakman
arXiv AI
Jun 16

Metacognitive Myopia in Large Language Models

arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.

By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.