Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper examines how large language models (LLMs) respond to different demographic cues—such as names—when users seek advice, focusing on race and gender in a U.S. context. It finds that using different cues for the same group leads to only partially overlapping changes in model responses, producing inconsistent conclusions about personalization and unstable bias metrics. The authors argue that LLMs react to linguistic signals tied to cues rather than to stable demographic categories, and they call for evaluations that use multiple cues and consider underlying mechanisms.
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2606. 08497v1 Announce Type: new Abstract: As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability.
arXiv:2508.08855v5 Announce Type: replace-cross Abstract: Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation st...
arXiv:2609.00491v1 Announce Type: new Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect larg...
arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.