From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale
Read the original on arXiv AI →The paper investigates covert dialect bias in large language models (LLMs) by analyzing how internal probability distributions associate different English varieties—Standard American English, African American Vernacular English, Nigerian Standard English, and Nigerian Pidgin—with housing-related adjectives. Using 260 meaning‑matched sentence quadruples and log‑probability scoring across ten open‑weight LLMs, the study finds that AAVE and NP are consistently linked to more negative adjectives than SAE, with NP experiencing the greatest penalty. The bias varies by context and stereotype cluster, and Nigerian Standard English shows a context‑dependent shift, being favored in formal tenant screening but penalized in more socially proximate scenarios.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.