arXiv Machine Learning By Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri, Raphael H. Heiberger, Steffen Staab

STEREODISCO: Discovering Stereotypicality in LLMs

Read the original on arXiv Machine Learning →

arXiv:2607. 27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 29

Localizing Persona Representations in LLMs

arXiv:2505. 24539v4 Announce Type: replace-cross Abstract: We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs).

By Celia Cintas, Miriam Rateike, Erik Miehling, Elizabeth Daly, Skyler Speakman