Hugging Face Trending Papers

Locating and Controlling Implicit Personalization in Large Language Models

Read the original on Hugging Face Trending Papers →

Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.