Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2603. 23485v2 Announce Type: replace-cross Abstract: Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses.
arXiv:2601. 05751v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade.
arXiv:2606. 07969v1 Announce Type: cross Abstract: Gender bias in AI-generated stories is a well-documented problem.
arXiv:2607. 28319v1 Announce Type: cross Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs).
arXiv:2608. 13328v1 Announce Type: cross Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally?
The paper examines how large language models (LLMs) respond to different demographic cues—such as names—when users seek advice, focusing on race and gender in a U.S. context. It finds that using different cues for the same group leads to only partially overlapping changes in model responses, producing inconsistent conclusions about personalization and unstable bias metrics. The authors argue that LLMs react to linguistic signals tied to cues rather than to stable demographic categories, and they call for evaluations that use multiple cues and consider underlying mechanisms.